Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 18 additions & 3 deletions docs/API.md
Original file line number Diff line number Diff line change
Expand Up @@ -1059,8 +1059,8 @@ The events worth grepping for have a section each below, and the index is a link
the index when you have a `type` off a log line and want to know what it means; read a section when
you want to know what the field it names is for and what it costs.

**The index is the whole log.** `src/` emits **121** event types and every one of them has a section
below — **115** sections, because a few cover a pair of events that are only read together. So a
**The index is the whole log.** `src/` emits **124** event types and every one of them has a section
below — **116** sections, because a few cover two or three events that are only read together. So a
`type` you cannot find here is not one the index skipped: it is a misread line, or a name `src/` no
longer emits.

Expand Down Expand Up @@ -1186,6 +1186,7 @@ emits fails it too.
| [`contribution_failed`](#contribution_failed) | The filing step threw, **after** `run_complete` |
| [`run_failed`](#run_failed) | The run threw, so there is **no document** |
| [`tagged_pdf` / `tagged_pdf_failed`](#tagged_pdf--tagged_pdf_failed) | A tagged PDF was made, or could not be |
| [`form_fields` / `page_fields` / `page_fields_missing`](#form_fields--page_fields--page_fields_missing) | The PDF's form field names were read, shown to a page, or not used |
| [`calibrate_call_failed`](#calibrate_call_failed) | One calibration verifier call threw — a tool's line, never a run's |

### `page_refit`
Expand Down Expand Up @@ -1955,7 +1956,7 @@ carried one. `where` is what makes the count attributable: the same character fr
from the correction pass and from a specialist are three facts about three different calls — and
`redrawn: true` is present when the reply was a page's [second draw](#page_redrawn), whose markup Iris
discarded, because a redraw makes two `extract` calls for one page. The same flag appears for the same
reason on `page_style_attributes`, `page_digit_groups` and `page_links`. It is
reason on `page_style_attributes`, `page_digit_groups`, `page_links` and `page_fields`. It is
written AFTER `agent_call`, so the reply on record in the round logs is still the model's own —
the census behind this row was a $0 regrade of logs already on disk, and a strip applied before
the log would have left no way to take that measurement or any future one.
Expand Down Expand Up @@ -4368,6 +4369,20 @@ A [tagged PDF](#get-a-tagged-pdf-optional) was made, or the tagger refused. `ms`
never logged. `tagged_pdf_failed` has the tagger's `code` and `error`. An `internal_error`'s
message is left out, because it may quote a value.

### `form_fields` / `page_fields` / `page_fields_missing`

With tagged PDFs on, the upload reads the PDF's form fields, and the page agent is asked to name
each control after its field, so the tagger can tag the field where it sits.

- `form_fields`: at upload, `fields` (how many were kept) and `unusable` (fields dropped: a name with
a quote, `'`, `<`, `>`, a backtick or a control character, a name over 100 characters, or no page
number), or the tagger's `error` code. Without them the run converts the same.
- `page_fields`: `image`, `fields` (how many were in the prompt), `dropped` (past the cap of 40) and,
on a second draw, `redrawn: true`.
- `page_fields_missing`: `image`, and `fields`, the names no control in the first pass's HTML has.
Only logged; no correction is made. A control a later correction removes is not reported here, but
the tagged PDF's report lists it as `field_not_in_html`.

### `calibrate_call_failed`

One verifier call in the calibration harness threw: `image` is the page, `defect` the seeded defect the
Expand Down
3 changes: 3 additions & 0 deletions src/pipeline/context.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ import type { Paths } from "../store/paths.ts";
import type { RunLog } from "../store/runlog.ts";
import type { IrisConfig } from "../config.ts";
import type { PdfLink } from "../util/pdf.ts";
import type { PdfField } from "../util/taggedPdf.ts";

export interface InputImage {
name: string; // filename, e.g. page-001.png
Expand All @@ -17,6 +18,8 @@ export interface InputImage {
// something other than an upload (the regression gate's fixture images) have no
// links to give it.
links?: PdfLink[];
// The source PDF's form fields on this page, when tagged PDFs are on (pipeline/fields.ts).
fields?: PdfField[];
}

// Everything a pipeline phase needs. Created once per run.
Expand Down
21 changes: 19 additions & 2 deletions src/pipeline/extraction.ts
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ import {
import { examplesForPrompt } from "./memory.ts";
import { altTexts, genericAltProblem, genericAlts } from "./alt.ts";
import { missingLinkProblem, missingLinks, pageLinkContext, unexpectedHrefs } from "./links.ts";
import { missingFields, pageFieldContext } from "./fields.ts";
import { duplicateIdProblem, duplicateIds, idAudit } from "./anchors.ts";
import { splitWordAudit, splitWordContradictions, splitWordProblem } from "./hyphens.ts";
import { STANDARD as STANDARD_AGENTS, isStandardType, logicalType } from "./contribute.ts";
Expand Down Expand Up @@ -3747,9 +3748,19 @@ async function renderPage(
...(redrawn ? { redrawn: true } : {}),
});
}
// And its form fields' names, which the image cannot show either (pipeline/fields.ts).
const fields = pageFieldContext(img.fields);
if (fields.shown.length) {
ctx.log.event("page_fields", {
image: img.name,
fields: fields.shown.length,
dropped: fields.dropped,
...(redrawn ? { redrawn: true } : {}),
});
}
const user =
`Convert this document page image (filename: ${img.name}, page ${img.order} of ${ctx.images.length}) ` +
`to accessible HTML.${links.section}${feedbackPreamble(ctx)}${priorSection}`;
`to accessible HTML.${links.section}${fields.section}${feedbackPreamble(ctx)}${priorSection}`;
const res = await ctx.router.complete(
PAGE_AGENT,
"vision",
Expand Down Expand Up @@ -4326,7 +4337,7 @@ async function correctPage(
`Omit it where you are acting on every problem. Return the page in "html" either way — ` +
`unchanged where you declined everything — because a reply with no "html" is a reply this run ` +
`cannot use, and every problem you did not decline is still to be fixed in the same reply.` +
`${pageLinkContext(img.links).section}`;
`${pageLinkContext(img.links).section}${pageFieldContext(img.fields).section}`;
const res = await ctx.router.complete(
PAGE_AGENT,
"vision",
Expand Down Expand Up @@ -5021,6 +5032,12 @@ async function extractPage(
if (missing.length) {
ctx.log.event("page_links_missing", { image: img.name, links: missing.map((l) => l.href) });
}
// Fields with no control named after them are only logged, not corrected: the first
// measurement is whether the names match at all (#483).
const unnamed = missingFields(img.fields, innerHtml);
if (unnamed.length) {
ctx.log.event("page_fields_missing", { image: img.name, fields: unnamed.map((f) => f.name) });
}

// And whether any image on the page was described with a placeholder instead of a
// description, checked here for the same reason and on the same terms: it has an exact answer,
Expand Down
97 changes: 97 additions & 0 deletions src/pipeline/fields.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,97 @@
import type { PdfField } from "../util/taggedPdf.ts";
import { decodeEntities } from "../util/html.ts";

// Naming a PDF's form controls after its own fields (#483).
//
// When tagged PDFs are on, iris-pdf tags each form field where its control sits in the
// HTML, and it finds the control by `name`. The page image does not show field names, so
// the upload reads them from the PDF (`iris-pdf fields`) and they are listed in the page
// agent's prompt, the way links.ts lists link targets. `missingFields` checks the reply.
// A field with no control is logged, not yet corrected.

export const MAX_FIELDS_PER_PAGE = 40;
export const MAX_NAME_CHARS = 100;

// Characters that would end the `name="…"` attribute the agent copies a name into, as for
// link hrefs (`UNSAFE_CHARS` in util/pdf.ts). A space is allowed: field names have them, and
// a space cannot leave a quoted value. A field refused here is tagged at the end of its
// page, as every field was before #483.
const UNSAFE_NAME = /["'<>`\u0000-\u001f\u007f]/;

export function usableFieldName(name: unknown): name is string {
return typeof name === "string" && name.length > 0 && name.length <= MAX_NAME_CHARS && !UNSAFE_NAME.test(name);
}

// Choice lists longer than this, or with longer entries, are left out of the prompt.
const MAX_OPTIONS = 10;
const MAX_OPTION_CHARS = 40;
const CHOICE_TYPES = new Set(["radio", "combobox", "listbox"]);

function describe(f: PdfField): string {
const parts: string[] = [];
// iris-pdf's own type word. Anything else is not printed.
if (/^[a-z]+$/.test(f.type ?? "")) parts.push(f.type);
const options = Array.isArray(f.options) ? f.options : [];
if (
CHOICE_TYPES.has(f.type) &&
options.length > 0 &&
options.length <= MAX_OPTIONS &&
options.every((o) => typeof o === "string" && o.length <= MAX_OPTION_CHARS)
) {
parts.push(`options: ${options.map((o) => JSON.stringify(o)).join(", ")}`);
}
return parts.length ? ` (${parts.join("; ")})` : "";
}

// The page's form fields, as the section of the page-agent prompt that carries them, plus
// what was dropped to bound it. Empty section when the page has none, so a deployment
// without iris-pdf sends exactly the prompt it sent before.
export function pageFieldContext(fields: PdfField[] = []): {
section: string;
shown: PdfField[];
dropped: number;
} {
if (fields.length === 0) return { section: "", shown: [], dropped: 0 };
const shown = fields.slice(0, MAX_FIELDS_PER_PAGE);
const dropped = fields.length - shown.length;
const list = shown
// JSON-quoted: the name is the PDF's, and a quote or newline in it must not end the quote.
.map((f, i) => `${i + 1}. ${JSON.stringify(f.name)}${describe(f)}`)
.join("\n");
const section =
`\n\n## Form fields on this page (from the source file's own form fields)\n` +
`The source file is a fillable form, and these are its fields on this page. Give each ` +
`field's control a name attribute holding the field's name EXACTLY as listed (the text ` +
`inside the quotes). The names are not labels: keep the label the page prints.\n\n` +
`${list}\n\n` +
(dropped > 0 ? `(…and ${dropped} more field${dropped === 1 ? "" : "s"} on this page.)\n\n` : "") +
`Do not add a control the image does not show, and do not invent names for controls that ` +
`are not listed. If you cannot tell which control a field belongs to, say so in the "log" ` +
`field.\n`;
return { section, shown, dropped };
}

// Every name attribute's value in a fragment of HTML. A scan, like links.ts's `hrefsIn`,
// and preceded by whitespace or a quote so `data-name=` does not count.
function namesIn(html: string): Set<string> {
const found = new Set<string>();
for (const m of html.matchAll(/(?<=[\s"'])name\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'>]+))/gi)) {
found.add(decodeEntities(m[1] ?? m[2] ?? m[3] ?? ""));
}
return found;
}

// The listed fields no control in the HTML is named after. Deduplicated by name: a radio
// group is one field with several controls.
export function missingFields(fields: PdfField[] = [], html: string): PdfField[] {
if (fields.length === 0) return [];
const present = namesIn(html);
const seen = new Set<string>();
const missing: PdfField[] = [];
for (const f of fields.slice(0, MAX_FIELDS_PER_PAGE)) {
if (present.has(f.name) || seen.has(f.name)) continue;
seen.add(f.name);
missing.push(f);
}
return missing;
}
18 changes: 13 additions & 5 deletions src/pipeline/orchestrator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -41,18 +41,19 @@ import { altTexts, genericAlts } from "./alt.ts";
import { markupReport } from "./markup.ts";
import type { Fragment } from "./fragment.ts";
import type { PdfLink } from "../util/pdf.ts";
import type { PdfField } from "../util/taggedPdf.ts";

// The link annotations the upload extracted from its PDFs, keyed by page order
// (see Paths.sessionLinks). Absent for a session of plain images, for a PDF with no
// links, and for any session created before links were extracted at all — all of
// which mean the same thing here, so a missing or unreadable file is no links rather
// than an error. Links are additive: without them a run produces the document it
// always produced.
function readLinks(paths: Paths, sessionId: string): Record<string, PdfLink[]> {
const path = paths.sessionLinks(sessionId);
// The same holds for form fields (fields.json, #483).
function readByOrder<T>(path: string): Record<string, T[]> {
if (!existsSync(path)) return {};
try {
const parsed = JSON.parse(readFileSync(path, "utf8")) as Record<string, PdfLink[]>;
const parsed = JSON.parse(readFileSync(path, "utf8")) as Record<string, T[]>;
return parsed && typeof parsed === "object" ? parsed : {};
} catch {
return {};
Expand All @@ -63,13 +64,20 @@ function readLinks(paths: Paths, sessionId: string): Record<string, PdfLink[]> {
// (which is significant — see docs/API.md) survives, independent of filename.
export function enumerateInputs(paths: Paths, sessionId: string): InputImage[] {
const dir = paths.sessionInput(sessionId);
const links = readLinks(paths, sessionId);
const links = readByOrder<PdfLink>(paths.sessionLinks(sessionId));
const fields = readByOrder<PdfField>(paths.sessionFields(sessionId));
return readdirSync(dir)
.filter((f) => f.includes("__"))
.map((f) => {
const [prefix, ...rest] = f.split("__");
const order = parseInt(prefix, 10);
return { order, name: rest.join("__"), path: join(dir, f), links: links[String(order)] ?? [] };
return {
order,
name: rest.join("__"),
path: join(dir, f),
links: links[String(order)] ?? [],
fields: fields[String(order)] ?? [],
};
})
.sort((a, b) => a.order - b.order);
}
Expand Down
39 changes: 36 additions & 3 deletions src/routes/sessions.ts
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,8 @@ import {
shrunkPageRejection,
} from "../providers/imageLimits.ts";
import { imageDimensions } from "../util/imageSize.ts";
import { readFields, tagPdf, taggedPdfCommand, tagTimeoutSeconds, TaggedPdfError } from "../util/taggedPdf.ts";
import { usableFieldName } from "../pipeline/fields.ts";
import { readFields, tagPdf, taggedPdfCommand, tagTimeoutSeconds, TaggedPdfError, type PdfField } from "../util/taggedPdf.ts";

// 50 MB is a memory bound, not the image limit. It stays well above what an image may
// be (see imageLimits.ts) because a PDF legitimately is: 25 pages of scans is a large
Expand Down Expand Up @@ -157,6 +158,35 @@ export function keepSourcePdf(
return true;
}

// The kept PDF's form fields, by page, so the page agent can name each control after its
// field (#483, pipeline/fields.ts). Written only when there are some. A failure is returned
// for the run log rather than raised: the run converts the same without them. A name that
// could end the attribute it is copied into is dropped and counted (`usableFieldName`).
export async function keepPdfFields(
command: string,
paths: Paths,
sessionId: string,
): Promise<{ fields: number; unusable: number } | { error: string }> {
let fields: PdfField[];
try {
fields = await readFields(command, paths.sessionSourcePdf(sessionId));
} catch (e) {
return { error: e instanceof TaggedPdfError ? e.code : "tagger_failed" };
}
if (!Array.isArray(fields)) return { error: "tagger_failed" };
const byPage: Record<string, PdfField[]> = {};
let unusable = 0;
for (const f of fields) {
if (!usableFieldName(f?.name) || !Number.isInteger(f.page)) {
unusable++;
continue;
}
(byPage[String(f.page)] ??= []).push(f);
}
if (Object.keys(byPage).length) writeFileSync(paths.sessionFields(sessionId), JSON.stringify(byPage, null, 2));
return { fields: fields.length - unusable, unusable };
}

export function sessionsRouter(cfg: IrisConfig, store: Store): Router {
const r = Router();
const paths = new Paths(cfg);
Expand Down Expand Up @@ -442,16 +472,19 @@ export function sessionsRouter(cfg: IrisConfig, store: Store): Router {
if (Object.keys(linksByOrder).length) {
writeFileSync(paths.sessionLinks(sessionId), JSON.stringify(linksByOrder, null, 2));
}
keepSourcePdf(cfg, paths, sessionId, files);
const formFields = keepSourcePdf(cfg, paths, sessionId, files)
? await keepPdfFields(taggedPdfCommand(cfg)!, paths, sessionId)
: null;
// A page the model reads at lower resolution than the rest is a fact about this
// document's output, so the session says so rather than the deployment's stdout: the
// owner of a document whose fold-out reads worse than its letter pages can find out
// why from the log they already have (`GET /v1/sessions/{id}/logs`). Best-effort for
// the same reason `run_queued` is — a run must not fail to start over its own log.
if (refits.length) {
if (refits.length || formFields) {
try {
const log = new RunLog(paths.sessionLog(sessionId));
for (const r of refits) log.event("page_refit", r);
if (formFields) log.event("form_fields", formFields);
} catch {
// ignore — observability must not block the run
}
Expand Down
9 changes: 7 additions & 2 deletions src/store/paths.ts
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,8 @@ import type { IrisConfig } from "../config.ts";
// final.json), history/ (the PRIOR output.html, snapshotted only when a
// feedback re-run is about to overwrite it — not the review loop's rounds),
// output.html, log.jsonl (the run log), lint.json, unresolved.md,
// agent-updates.md, links.json, source-name.txt, source.pdf (only with
// tagged PDFs on)
// agent-updates.md, links.json, source-name.txt, source.pdf and
// fields.json (only with tagged PDFs on)
// fixtures/<agent>/ and memory/<agent>.json — keyed by agent, shared by every session
// tmp/<id>/ one run's scratch. `tmp/<id>/agents/` holds agents that session BUILT,
// which `loadAgent` prefers over the library for the rest of it.
Expand Down Expand Up @@ -68,6 +68,11 @@ export class Paths {
sessionLinks(id: string): string {
return join(this.sessionDir(id), "links.json");
}
// The source PDF's form fields by page order, like links.json (#483). Written only when
// there are some.
sessionFields(id: string): string {
return join(this.sessionDir(id), "fields.json");
}
// The uploaded PDF, kept only when tagged PDFs are on and the upload was one PDF.
// It is what `iris-pdf` tags (util/taggedPdf.ts).
sessionSourcePdf(id: string): string {
Expand Down
3 changes: 3 additions & 0 deletions test/fixtures/fake-iris-pdf.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,9 @@ if (command === "fields") {
console.log(JSON.stringify([
{ name: "applicant.name", type: "text", page: 1, options: [], required: true, readonly: false, maxlen: 40, editable: false, multiSelect: false },
{ name: "applicant.consent", type: "checkbox", page: 1, options: ["Yes"], required: false, readonly: false, maxlen: null, editable: false, multiSelect: false },
...(pdf.includes("UNSAFE")
? [`x" onfocus="alert(1)`, "y".repeat(101)].map((name) => ({ name, type: "text", page: 1, options: [], required: false, readonly: false, maxlen: null, editable: false, multiSelect: false }))
: []),
]));
} else if (command === "tag") {
const values = JSON.parse(readFileSync(a.values, "utf8"));
Expand Down
Loading
Loading