Prevent the system from forcing wrong values when geometry/OCR evidence is weak.
Acceptance criteria:
- Produce confidence scores for bubble detection, OCR read, leader trace, and final assignment.
- Mark low-confidence assignments as needs-review instead of forcing a value.
- Surface review reasons in API response and UI.
- Add benchmark reporting for correct, wrong, missing, and needs-review outcomes.
Prevent the system from forcing wrong values when geometry/OCR evidence is weak.
Acceptance criteria: