Iris Performance Optimizer Agent here.
The gap
The page corpus has no page whose meaning lives in a graphic. Not one chart, diagram, data visualisation, org chart, flow chart, map or annotated figure where the information is in the picture and nowhere else on the page.
Two pages are named as if they were exactly that, and neither is:
| page id |
what is actually on it |
meta-chart-p20 |
prose over a decorative sky photograph. No chart. |
meta-infographic-p30 |
prose plus a pull quote beside a decorative sunset. No infographic. |
On both of those, alt="" is the correct markup under WCAG 1.1.1, because the images carry no information.
Why it matters
It leaves the bar unable to test the single capability that most sharply separates the candidate approaches from each other.
Describing a graphic is the one job in page extraction that a vision model can do and OCR structurally cannot. Textract will report FIGURE and a bounding box; it will never tell you that the bars show enrolment falling 12% between 2019 and 2023. If the sprint's answer for the extraction step ends up being "OCR for structure plus a cheap model", the question of which model is largely a question about graphic description — and right now the bench cannot see that difference at all. Every approach in the ranking scores identically on it, including the ones that emit nothing.
structureDefects never looks at images, and axe treats alt="" as valid because it means "decorative". So an extractor that emits <figure><img alt=""></figure> where a chart should be described scores clean today.
How this surfaced
I built the check for it (altcheck.py), ran it, and got a clean-looking split: every no-model approach described nothing, 12 of 15 model answers described something. I was about to promote it to a gate, and I had already run the gate-before-you-fail check (altgate.py) showing it would newly fail only one extractor.
Then I opened the two PNGs. The gate would have failed the OCR extractors for producing the correct answer, and some of the models it rewarded are writing descriptions of scenery that carries no information — which is itself a WCAG defect, not a virtue.
I had taken the page IDs as facts. They are filenames somebody typed. Both scripts now carry the reversal in their header comments; neither gate exists.
What would close it
Three or four pages added to the corpus where the page's meaning is genuinely in a graphic, each with ground truth for what the graphic says, not just for the words printed around it:
- a bar or line chart carrying a trend that appears nowhere in the body text
- a data table rendered as an image (no text layer for it at all)
- a process or flow diagram with labelled steps and directional edges
- a map or org chart with spatial relationships that matter
The hard part is the ground truth, and it is a different kind of ground truth from the rest of the corpus. Everywhere else, truth is the PDF's own text layer and scoring is word overlap. A graphic has no text layer to compare against, and "is this description adequate" is not a word-overlap question. Two options, and I don't think this should be decided by whoever writes the code:
- Hand-written reference descriptions plus a keyed fact list — score on whether the required facts (direction, magnitude, endpoints, labels) appear. Objective and rerunnable; expensive to author, and it bakes in one author's judgement of what matters.
- A rubric applied by a judge model — cheaper to extend, but it makes the bar depend on a model's opinion, which this bench has deliberately avoided everywhere else. Model consensus was explicitly rejected earlier in the sprint as a self-consistency trap; a judge is a milder version of the same problem.
I lean to (1) for a small number of pages, precisely because it is falsifiable in the direction that matters: someone can disagree with a keyed fact and check.
Scope note
This is a corpus gap, not a bar gap, and it is not a blocker for the extraction-step recommendation in EqualifyEverything/equalify-iris#246 — the models that lead that ranking lead it on word capture and structure, which are measured. It is a stated limitation on any claim of the form "OCR is good enough for extraction", and it should be closed before that claim gets made.
Iris Performance Optimizer Agent here.
The gap
The page corpus has no page whose meaning lives in a graphic. Not one chart, diagram, data visualisation, org chart, flow chart, map or annotated figure where the information is in the picture and nowhere else on the page.
Two pages are named as if they were exactly that, and neither is:
meta-chart-p20meta-infographic-p30On both of those,
alt=""is the correct markup under WCAG 1.1.1, because the images carry no information.Why it matters
It leaves the bar unable to test the single capability that most sharply separates the candidate approaches from each other.
Describing a graphic is the one job in page extraction that a vision model can do and OCR structurally cannot. Textract will report
FIGUREand a bounding box; it will never tell you that the bars show enrolment falling 12% between 2019 and 2023. If the sprint's answer for the extraction step ends up being "OCR for structure plus a cheap model", the question of which model is largely a question about graphic description — and right now the bench cannot see that difference at all. Every approach in the ranking scores identically on it, including the ones that emit nothing.structureDefectsnever looks at images, and axe treatsalt=""as valid because it means "decorative". So an extractor that emits<figure><img alt=""></figure>where a chart should be described scores clean today.How this surfaced
I built the check for it (
altcheck.py), ran it, and got a clean-looking split: every no-model approach described nothing, 12 of 15 model answers described something. I was about to promote it to a gate, and I had already run the gate-before-you-fail check (altgate.py) showing it would newly fail only one extractor.Then I opened the two PNGs. The gate would have failed the OCR extractors for producing the correct answer, and some of the models it rewarded are writing descriptions of scenery that carries no information — which is itself a WCAG defect, not a virtue.
I had taken the page IDs as facts. They are filenames somebody typed. Both scripts now carry the reversal in their header comments; neither gate exists.
What would close it
Three or four pages added to the corpus where the page's meaning is genuinely in a graphic, each with ground truth for what the graphic says, not just for the words printed around it:
The hard part is the ground truth, and it is a different kind of ground truth from the rest of the corpus. Everywhere else, truth is the PDF's own text layer and scoring is word overlap. A graphic has no text layer to compare against, and "is this description adequate" is not a word-overlap question. Two options, and I don't think this should be decided by whoever writes the code:
I lean to (1) for a small number of pages, precisely because it is falsifiable in the direction that matters: someone can disagree with a keyed fact and check.
Scope note
This is a corpus gap, not a bar gap, and it is not a blocker for the extraction-step recommendation in EqualifyEverything/equalify-iris#246 — the models that lead that ranking lead it on word capture and structure, which are measured. It is a stated limitation on any claim of the form "OCR is good enough for extraction", and it should be closed before that claim gets made.