Skip to content

Corpus has no page whose meaning lives in a graphic, so the bar cannot test the one thing OCR cannot do #1

Description

@bbertucc

Iris Performance Optimizer Agent here.

The gap

The page corpus has no page whose meaning lives in a graphic. Not one chart, diagram, data visualisation, org chart, flow chart, map or annotated figure where the information is in the picture and nowhere else on the page.

Two pages are named as if they were exactly that, and neither is:

page id what is actually on it
meta-chart-p20 prose over a decorative sky photograph. No chart.
meta-infographic-p30 prose plus a pull quote beside a decorative sunset. No infographic.

On both of those, alt="" is the correct markup under WCAG 1.1.1, because the images carry no information.

Why it matters

It leaves the bar unable to test the single capability that most sharply separates the candidate approaches from each other.

Describing a graphic is the one job in page extraction that a vision model can do and OCR structurally cannot. Textract will report FIGURE and a bounding box; it will never tell you that the bars show enrolment falling 12% between 2019 and 2023. If the sprint's answer for the extraction step ends up being "OCR for structure plus a cheap model", the question of which model is largely a question about graphic description — and right now the bench cannot see that difference at all. Every approach in the ranking scores identically on it, including the ones that emit nothing.

structureDefects never looks at images, and axe treats alt="" as valid because it means "decorative". So an extractor that emits <figure><img alt=""></figure> where a chart should be described scores clean today.

How this surfaced

I built the check for it (altcheck.py), ran it, and got a clean-looking split: every no-model approach described nothing, 12 of 15 model answers described something. I was about to promote it to a gate, and I had already run the gate-before-you-fail check (altgate.py) showing it would newly fail only one extractor.

Then I opened the two PNGs. The gate would have failed the OCR extractors for producing the correct answer, and some of the models it rewarded are writing descriptions of scenery that carries no information — which is itself a WCAG defect, not a virtue.

I had taken the page IDs as facts. They are filenames somebody typed. Both scripts now carry the reversal in their header comments; neither gate exists.

What would close it

Three or four pages added to the corpus where the page's meaning is genuinely in a graphic, each with ground truth for what the graphic says, not just for the words printed around it:

  • a bar or line chart carrying a trend that appears nowhere in the body text
  • a data table rendered as an image (no text layer for it at all)
  • a process or flow diagram with labelled steps and directional edges
  • a map or org chart with spatial relationships that matter

The hard part is the ground truth, and it is a different kind of ground truth from the rest of the corpus. Everywhere else, truth is the PDF's own text layer and scoring is word overlap. A graphic has no text layer to compare against, and "is this description adequate" is not a word-overlap question. Two options, and I don't think this should be decided by whoever writes the code:

  1. Hand-written reference descriptions plus a keyed fact list — score on whether the required facts (direction, magnitude, endpoints, labels) appear. Objective and rerunnable; expensive to author, and it bakes in one author's judgement of what matters.
  2. A rubric applied by a judge model — cheaper to extend, but it makes the bar depend on a model's opinion, which this bench has deliberately avoided everywhere else. Model consensus was explicitly rejected earlier in the sprint as a self-consistency trap; a judge is a milder version of the same problem.

I lean to (1) for a small number of pages, precisely because it is falsifiable in the direction that matters: someone can disagree with a keyed fact and check.

Scope note

This is a corpus gap, not a bar gap, and it is not a blocker for the extraction-step recommendation in EqualifyEverything/equalify-iris#246 — the models that lead that ranking lead it on word capture and structure, which are measured. It is a stated limitation on any claim of the form "OCR is good enough for extraction", and it should be closed before that claim gets made.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions