Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Use accessible HTML to tag and update a PDF.

[Iris](https://github.com/EqualifyEverything/equalify-iris) turns page images into accessible HTML. This tool takes that HTML and the original PDF, and gives back **the same PDF, tagged**: a structure tree a screen reader can follow, with form fields filled in if you give it values. The page looks exactly as it did.

Tagging runs offline and makes no network or model calls. An optional `review` asks a Claude model to check the result.
Tagging runs offline and makes no network or model calls. An optional `review` asks a model to check the result.

## Install

Expand Down Expand Up @@ -79,9 +79,9 @@ Then two checks run, and if either fails nothing is written (exit 2):

## Review

`review` checks what `check` cannot: whether the tags say what the page says. For each page it sends a Claude model the page image and what a screen reader gets from the page: the structure, text, alt text, link targets and field names. It reports missing content, wrong reading order, wrong element types or heading levels, tables, alt text, link text, field names and language. It prints one finding per line, writes them to `--report` as JSON with the tokens used and an estimated cost, and exits 0. A page the model could not review is reported with its error, the other pages are kept, and the exit is 1. Nothing in the PDF is changed.
`review` checks what `check` cannot: whether the tags say what the page says. For each page it sends a model the page image and what a screen reader gets from the page: the structure, text, alt text, link targets and field names. It reports missing content, wrong reading order, wrong element types or heading levels, tables, alt text, link text, field names and language. It prints one finding per line, writes them to `--report` as JSON with the tokens used and an estimated cost, and exits 0. A page the model could not review is reported with its error, the other pages are kept, and the exit is 1. Nothing in the PDF is changed.

**It sends page images and text to the model provider.** With `ANTHROPIC_API_KEY` set, it uses the Anthropic API. Otherwise it uses Amazon Bedrock through the AWS CLI, with your AWS credentials and region. Choose with `--provider anthropic|bedrock` and `--model <id>`. The default model is Opus 5.5, at about US$0.03 a page. See [docs/models.md](docs/models.md) for the models compared and their costs.
**It sends page images and text to the model provider.** With `ANTHROPIC_API_KEY` set, it uses the Anthropic API. Otherwise it uses Amazon Bedrock through the AWS CLI, with your AWS credentials and region. Choose with `--provider anthropic|bedrock` and `--model <id>`. The default model is Sonnet 5, at about US$0.02 a page; on Bedrock, `--model` takes any model that reads images and calls tools. See [docs/models.md](docs/models.md) for the models compared and their costs.

## PDF/UA

Expand Down
74 changes: 52 additions & 22 deletions docs/models.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# Models and costs

Where this project uses a Claude model, which one, and what it costs. `tag`, `fields` and `check` use no model and cost nothing to run.
Where this project uses a model, which one, and what it costs. `tag`, `fields` and `check` use no model and cost nothing to run.

List prices, USD per million tokens (September 2026):
Claude list prices, USD per million tokens (September 2026):

| Model | Input | Output |
|---|---|---|
Expand All @@ -12,26 +12,56 @@ List prices, USD per million tokens (September 2026):
| Opus 5 | 5 | 25 |
| Fable 5.1 | 10 | 50 |

These apply to the Anthropic API and to Bedrock's `global.` inference profiles. Bedrock's regional profiles (`us.`, `eu.`, …) cost 10% more. Our AWS organization allows only the `us.` profiles, so the figures below include that 10%.

## `iris-pdf review`: Opus 5.5

**Use Opus 5.5, the default.** If cost matters more than precision, `--model us.anthropic.claude-sonnet-5` (or `claude-sonnet-5`) costs about two thirds as much. Haiku 4.5 is not recommended.

Measured on 2026-09-24 on Bedrock `us.` profiles, with the prompt in `src/review/review.ts`:

| Model | Seeded defects found (of 8) | Findings on 3 clean pages | 25-page scanned report (ACIR): cost, per page, time |
|---|---|---|---|
| Opus 5.5 | 8 | 0 | $0.72, $0.029, 54 s |
| Sonnet 5 | 8 | 0 | $0.47, $0.019, 82 s |
| Haiku 4.5 | 8 | 4 | $0.14, $0.005, 35 s |

- **The seeded defects** were fixture pages tagged from altered HTML: a heading tagged as a paragraph, a skipped heading level, swapped columns, a list and a table tagged as paragraphs, generic alt text, and Chinese text in a document declared English.
- **On real documents** both Opus and Sonnet found real problems: links with no destination, captions that aren't on the page, and OCR noise tagged as text.
- Sonnet made more mistakes. Before the prompt said so, it misread `Art` (Article) as an artifact. It also writes about twice as many output tokens, which is why it is only about a third cheaper.
- Haiku reported running headers as missing content despite being told they are artifacts. Once it answered without calling the findings tool.
- **Fable 5.1** was not available on Bedrock to this account when measured. At 2.5 times Opus 5.5's price, it isn't needed for this task.
- **Tokens.** A page is about 4,000 input tokens (the image is most of it) and 300–850 output tokens.
These apply to the Anthropic API and to Bedrock's `global.` inference profiles. Bedrock's regional profiles (`us.`, `eu.`, …) cost 10% more. Our AWS organization allows only the `us.` profiles, so the figures below include that 10%. Other vendors' models on Bedrock are billed at AWS's published rates.

## `iris-pdf review`: Sonnet 5

**Use Sonnet 5, the default.** It found as much as any model we measured, at the lowest price among the best. On a budget, **`--model us.openai.gpt-5.6-luna`** (Bedrock only) costs about a fifteenth as much and misses a little more.

On Bedrock, `review` uses the Converse API, so `--model` takes any Bedrock model that reads images and calls tools.

### How it was measured

On 2026-09-24, through Bedrock `us.` profiles, with the prompt in `src/review/review.ts`. The method follows equalify-iris's verifier calibration: count both what a model catches and what it invents.

- **Corpus: 26 one-page PDFs.**
- 8 real pages from a 1962 scanned report (ACIR), tagged from equalify-iris's HTML.
- 8 of this repo's fixtures.
- 10 damaged copies with one defect each: a heading as a paragraph, a skipped heading level, two paragraphs swapped, a table as a paragraph, header cells as data cells, generic alt text, wrong alt text, a figure left untagged, a document language that is wrong, and a paragraph that is not on the page.
- **Defects caught:** a finding of the right kind that names the defect.
- **Real problems found:** the clean copies turned out to hold 15 real problems the tagger left. They are stray OCR fragments, dot leaders tagged as text, a title tagged twice, table cells pointing at a missing header, a link with no text, a figure with no image, and an unnamed button. These were checked by hand against the page images.
- **False positives:** findings on clean copies that are neither of those. An example is a running header reported as missing, though artifacts are excluded by design.
- **Draws:** three for the leaders and two for the rest. First, 26 models from 10 vendors were screened with one draw.

| Model | Defects caught | Real problems found | False positives per page | Per 100 pages |
|---|---|---|---|---|
| Kimi K3 | 30/30 | 43/45 | 0.02 | $3.31 |
| Opus 5.5 | 30/30 | 41/45 | 0.04 | $2.62 |
| **Sonnet 5** | 30/30 | 40/45 | 0.04 | **$1.93** |
| GPT-5.6 sol | 30/30 | 39/45 | 0 | not published |
| GPT-6 astra | 30/30 | 38/45 | 0 | not published |
| GPT-5.5 | 20/20 | 27/30 | 0 | not published |
| GPT-6 sol | 20/20 | 23/30 | 0 | not published |
| GPT-5.4 | 20/20 | 24/30 | 0.59 | $1.10 |
| GPT-5.6 terra | 20/20 | 23/30 | 0.16 | $1.03 |
| GPT-6 luna | 20/20 | 18/30 | 0.31 | not published |
| **GPT-5.6 luna** | 28/30 | 38/45 | 0.15 | **$0.13** |
| Haiku 4.5 | 19/20 | 22/30 | 0.63 | $0.56 |
| Kimi K2.5 | 18/20 | 25/30 | 0.63 | $0.29 |
| Mistral Large 3 | 19/20 | 19/30 | 1.59 | $0.24 |

- **The leaders tie within the noise.** Kimi K3, Opus 5.5, Sonnet 5, GPT-5.6 sol, GPT-6 astra and GPT-5.5 all caught every defect and invented almost nothing. Sonnet 5 is the cheapest of them with a published price.
- **GPT-5.5, GPT-5.6 sol and the GPT-6 models** have no rate in AWS's price list, so their cost is unknown. Measure them again once they are priced.
- **GPT-5.6 luna** missed a heading tagged as a paragraph and header cells tagged as data cells, once each. It sometimes reports running headers. Its rate is published only for us-gov-west-1, and GovCloud rates tend to be higher, so treat its cost as an upper bound.
- **Screened out after one draw:**
- Llama 4 Maverick, Pixtral Large, Ministral 14B, Qwen3 VL 235B and Nova Pro caught 6 to 9 of 10 defects, with up to 2.4 false positives a page.
- Grok 4.6 failed 9 of 27 pages, mostly by running past 4,096 output tokens.
- Nova 2 Lite caught 1 defect; Llama 4 Scout caught none.
- Gemma 3 and Nemotron Nano 2 VL never called the tool.
- Nova Premier has reached end of life.
- Fable 5.1 was unavailable on Bedrock to this account.
- **The corpus is small**, and the leaders caught every defect in it. Widen it before choosing between them. The corpus and harness are not in this repo yet.
- **Tokens.** Averaged over the corpus, Sonnet 5 uses about 3,800 input and 1,000 output tokens a page; GPT-5.6 luna, about 1,800 and 500.

The review sends each page's image and its text to the model provider, so do not use it on documents that must not leave your machine.

Expand Down
62 changes: 42 additions & 20 deletions src/review/review.ts
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
// An optional AI review of a tagged PDF. Each page's image and its
// screen-reader outline go to a Claude model, which reports what a blind
// screen-reader outline go to a model, which reports what a blind
// reader would miss or get wrong. It reports; it changes nothing.
// This sends page images and text to the model provider.
import * as mupdf from "mupdf";
import { execFile } from "node:child_process";
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { openPdf, MAX_PAGES } from "../pdf/document.ts";
Expand All @@ -13,7 +13,7 @@ import { pageOutline, headingsBefore } from "./outline.ts";
import { IrisPdfError, EXIT, VERSION } from "../report.ts";

export type Provider = "anthropic" | "bedrock";
export const DEFAULT_MODEL: Record<Provider, string> = { anthropic: "claude-opus-5-5", bedrock: "us.anthropic.claude-opus-5-5" }; // see docs/models.md
export const DEFAULT_MODEL: Record<Provider, string> = { anthropic: "claude-sonnet-5", bedrock: "us.anthropic.claude-sonnet-5" }; // see docs/models.md

const KINDS = ["missing_content", "reading_order", "structure", "table", "alt_text", "link", "form", "language", "other"] as const;
export type Finding = { kind: (typeof KINDS)[number]; severity: "error" | "warning"; element: string; detail: string };
Expand All @@ -30,11 +30,13 @@ export type ReviewReport = {
export type Send = (body: Record<string, unknown>) => Promise<unknown>;
export type ReviewOptions = { provider?: Provider; model?: string; password?: string; send?: Send; concurrency?: number };

// USD per million input and output tokens, from Anthropic's list prices (September 2026).
// Bedrock's regional profiles (us., eu., …) cost 10% more; global. ones do not.
const PRICES: [RegExp, number, number][] = [
[/fable-5-1/, 10, 50], [/opus-5-5/, 4, 20], [/opus-5(?!-\d)/, 5, 25],
[/sonnet-5/, 2, 10], [/sonnet-4-6/, 3, 15], [/haiku-4-5/, 1, 5],
// USD per million input and output tokens (September 2026). Claude: Anthropic's list prices;
// Bedrock's regional profiles (us., eu., …) cost 10% more, global. ones do not.
// Others: the AWS price list for Bedrock, us-east-2 (GPT-5.6 luna: us-gov-west-1, the only region listed).
const PRICES: [RegExp, number, number, claude?: true][] = [
[/fable-5-1/, 10, 50, true], [/opus-5-5/, 4, 20, true], [/opus-5(?!-\d)/, 5, 25, true],
[/sonnet-5/, 2, 10, true], [/sonnet-4-6/, 3, 15, true], [/haiku-4-5/, 1, 5, true],
[/gpt-5\.6-luna/, 0.264, 1.584], [/kimi-k3/, 3.3, 16.5], [/mistral-large-3/, 0.5, 1.5],
];

const TOOL = "report_findings";
Expand Down Expand Up @@ -118,10 +120,12 @@ export async function review(pdf: Uint8Array, opts: ReviewOptions = {}): Promise
let res = await ask(req);
// With tool_choice auto a model can answer in text instead; ask once more.
// Only its text is kept: a tool_use turn would need a tool_result.
// Roles must alternate (Converse enforces it), so with no text the nudge joins the first turn.
if (!reported(res) && (res as Reply).stop_reason !== "max_tokens") {
const said = ((res as Reply).content ?? []).filter((c) => c.type === "text" && c.text);
const messages = [...(req.messages as unknown[]), ...(said.length ? [{ role: "assistant", content: said }] : []),
{ role: "user", content: `Answer by calling ${TOOL}.` }];
const nudge = `Answer by calling ${TOOL}.`, [first] = req.messages as { role: string; content: unknown[] }[];
const messages = said.length ? [first, { role: "assistant", content: said }, { role: "user", content: nudge }]
: [{ ...first, content: [...first.content, { type: "text", text: nudge }] }];
res = await ask({ ...req, messages });
}
report.pages[i] = { page: i + 1, findings: findings(res, i + 1) };
Expand Down Expand Up @@ -150,7 +154,7 @@ function image(page: mupdf.PDFPage): string {
return Buffer.from(page.toPixmap(mupdf.Matrix.scale(s, s), mupdf.ColorSpace.DeviceRGB, false).asPNG()).toString("base64");
}

// tool_choice stays auto: Opus 5.5 rejects a forced tool.
// tool_choice stays auto: some models, Opus 5.5 among them, reject a forced tool.
function request(model: string, png: string, reader: string): Record<string, unknown> {
return {
model, max_tokens: 4096, system: SYSTEM, tools: TOOLS,
Expand Down Expand Up @@ -192,7 +196,7 @@ export const plain = (s: string) => s.replace(/[\u0000-\u001f\u007f-\u009f]/g, "
export function cost(provider: Provider, model: string, usage: ReviewReport["usage"]): number | null {
const price = PRICES.find(([re]) => re.test(model));
if (!price) return null;
const regional = provider === "bedrock" && !model.startsWith("global.") ? 1.1 : 1;
const regional = price[3] && provider === "bedrock" && !model.startsWith("global.") ? 1.1 : 1;
return Math.round((usage.inputTokens * price[1] + usage.outputTokens * price[2]) * regional) / 1e6;
}

Expand All @@ -214,21 +218,39 @@ async function anthropic(body: Record<string, unknown>): Promise<unknown> {
}
}

// Bedrock's Converse API takes images and tools the same way for every vendor's model.
export function toConverse(body: Record<string, any>): Record<string, unknown> {
const block = (b: { type: string; text?: string; source?: { data: string } }) =>
b.type === "image" ? { image: { format: "png", source: { bytes: b.source!.data } } } : { text: b.text ?? "" };
return {
modelId: body.model,
system: [{ text: body.system }],
messages: body.messages.map((m: { role: string; content: string | { type: string }[] }) =>
({ role: m.role, content: (typeof m.content === "string" ? [{ type: "text", text: m.content }] : m.content).map(block) })),
toolConfig: { tools: body.tools.map((t: typeof TOOLS[0]) => ({ toolSpec: { name: t.name, description: t.description, inputSchema: { json: t.input_schema } } })) },
inferenceConfig: { maxTokens: body.max_tokens },
};
}

export function fromConverse(r: any): Reply {
const content = (r.output?.message?.content ?? []).flatMap((c: any) =>
c.toolUse ? [{ type: "tool_use", name: c.toolUse.name, input: c.toolUse.input }] : typeof c.text === "string" ? [{ type: "text", text: c.text }] : []);
return { content, stop_reason: r.stopReason, usage: { input_tokens: r.usage?.inputTokens ?? 0, output_tokens: r.usage?.outputTokens ?? 0 } };
}

// Through the AWS CLI, which brings the user's credentials and region, and retries.
async function bedrock(body: Record<string, unknown>): Promise<unknown> {
const { model, ...rest } = body;
const dir = mkdtempSync(join(tmpdir(), "iris-pdf-review-"));
try {
writeFileSync(join(dir, "in.json"), JSON.stringify({ anthropic_version: "bedrock-2023-05-31", ...rest }));
await new Promise<void>((resolve, reject) => execFile("aws", [
"bedrock-runtime", "invoke-model", "--model-id", String(model), "--body", `fileb://${join(dir, "in.json")}`,
"--content-type", "application/json", "--accept", "application/json", join(dir, "out.json"),
], { timeout: TIMEOUT_MS }, (err, _out, stderr) => {
if (!err) return resolve();
writeFileSync(join(dir, "in.json"), JSON.stringify(toConverse(body)));
const out = await new Promise<string>((resolve, reject) => execFile("aws", [
"bedrock-runtime", "converse", "--cli-input-json", `file://${join(dir, "in.json")}`, "--output", "json",
], { timeout: TIMEOUT_MS, maxBuffer: 1 << 26 }, (err, stdout, stderr) => {
if (!err) return resolve(stdout);
if ((err as NodeJS.ErrnoException).code === "ENOENT") return reject(new IrisPdfError("no_credentials", "The Bedrock provider needs the AWS CLI. Install it, or set ANTHROPIC_API_KEY.", EXIT.badInput));
reject(new IrisPdfError("review_failed", `Bedrock: ${stderr.trim() || err.message}`));
}));
return JSON.parse(readFileSync(join(dir, "out.json"), "utf8"));
return fromConverse(JSON.parse(out));
} finally {
rmSync(dir, { recursive: true, force: true });
}
Expand Down
Loading
Loading