Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 37 additions & 6 deletions docs/API.md
Original file line number Diff line number Diff line change
Expand Up @@ -663,10 +663,18 @@ base64**, which is **3.75 MB on disk**, on Amazon Bedrock. Ask the deployment in
assuming, since it moves with the configured model and provider:
[`GET /v1/limits`](#upload-limits-unauthenticated). A **PDF** is not measured against that
limit — the file you send is not what reaches the model, since Iris rasterizes its pages at its
own resolution — but each *rendered page* is, and a page over it fails with a `400` naming the
page and the PDF. That happens with large-format pages: rasterizing at a fixed DPI means the
page image scales with the physical page, so a letter page renders well inside the limit and an
ARCH-D drawing does not.
own resolution — but each *rendered page* is. That matters for large-format pages: rasterizing
at a fixed DPI means the page image scales with the physical page, so a letter page renders well
inside the limit and an ARCH-D drawing does not.

A page that renders too large is **rendered again, smaller**, rather than failing the upload —
those pixels are Iris's choice and not yours. The second render fits the long edge the model
reads where that is documented (so nothing it would have looked at is given up), and the
`max_dimension_px` ceiling where Iris has no published limits for the configured model (so only
pixels that could not have been sent at all are given up). The page appears in the session's run
log as [`page_refit`](#page_refit). A page that is *still* over after that — a dense photographic
or halftoned scan is over on bytes, not on size — fails with a `400` naming the page, the PDF,
and the size it was retried at.

Pixel dimensions are mostly **not** a limit worth planning around, and this is the common
misdiagnosis: a large-but-light image converts fine, while a small-but-heavy photo is what
Expand Down Expand Up @@ -1051,8 +1059,8 @@ The events worth grepping for have a section each below, and the index is a link
the index when you have a `type` off a log line and want to know what it means; read a section when
you want to know what the field it names is for and what it costs.

**The index is the whole log.** `src/` emits **120** event types and every one of them has a section
below — **114** sections, because a few cover a pair of events that are only read together. So a
**The index is the whole log.** `src/` emits **121** event types and every one of them has a section
below — **115** sections, because a few cover a pair of events that are only read together. So a
`type` you cannot find here is not one the index skipped: it is a misread line, or a name `src/` no
longer emits.

Expand All @@ -1064,6 +1072,7 @@ emits fails it too.

| `type` | What it records |
| --- | --- |
| [`page_refit`](#page_refit) | A page too large to send was rendered again, smaller, instead of refusing the document |
| [`run_queued` / `run_dequeued`](#run_queued--run_dequeued) | The run's wait for a concurrency slot |
| [`run_start`](#run_start) | The run's opening line: how many source pages, and which of the three paths it took |
| [`phase`](#phase) | The pipeline entered a phase |
Expand Down Expand Up @@ -1179,6 +1188,28 @@ emits fails it too.
| [`tagged_pdf` / `tagged_pdf_failed`](#tagged_pdf--tagged_pdf_failed) | A tagged PDF was made, or could not be |
| [`calibrate_call_failed`](#calibrate_call_failed) | One calibration verifier call threw — a tool's line, never a run's |

### `page_refit`

One page of an uploaded PDF was rendered a second time, smaller, because the first render came out too
large to send: `pdf` and `page` name it (the PDF's own page number, 1-based), `from` is the size the
first render produced (`8334x8334`), and `long_edge_px` is what the second one was fitted to. One line
per page refitted, so a document of fold-outs writes several.

This is the only event the upload route writes, and it is written before the run is queued —
rasterizing is what measures a page, so the pages have to be too big before there is a session to
record it against. A log that opens with `page_refit` rather than
[`run_queued`](#run_queued--run_dequeued) is therefore reading correctly.

The line exists because the page it names is read at a lower resolution than the rest of the document,
and nothing else in the log would say so. Rasterizing runs at a fixed DPI, so pixel count follows the
physical page: a letter page lands near 1275x1650 and a 55-inch drawing past 8000 px, above which the
vision model errors instead of downscaling. Those pixels are Iris's choice, not the caller's, so Iris
renders the page again instead of refusing the upload — nothing is cropped, only resolution is given
up. When a fold-out's findings read thinner than the letter pages around it, this is the line that
explains why, and re-exporting that page at a smaller trim size is the fix that gets it back. A page
still over the limit after the second render fails the upload instead, with a `400` that names the size
it was retried at.

### `run_queued` / `run_dequeued`

The run's wait for a concurrency slot: how busy the queue was when it was admitted (`running` of
Expand Down
69 changes: 69 additions & 0 deletions src/providers/imageLimits.ts
Original file line number Diff line number Diff line change
Expand Up @@ -648,6 +648,75 @@ export function rasterizedPageRejection(
return null;
}

// How small to render a rasterized page that does not fit, or null when rendering it
// again cannot help. Read only after `rasterizedPageRejection` has said there is a
// problem; what it answers is whether Iris can fix that problem itself.
//
// It can, because it chose the pixels. A page over either limit at util/pdf.ts's DPI is
// a page whose ink Iris rendered too big — the caller uploaded a PDF, and the remedy the
// rejection offers them (re-export at a smaller page size, re-save the pages as JPEGs)
// asks them to do by hand what the renderer can do exactly. Refusing the document was
// the honest thing to do while the only alternative was failing inside a model call four
// minutes later; it is not the honest thing to do when a second render fits.
//
// WHICH pixels may be given up is the only real question, and the answer differs by
// basis, which is why it is decided here rather than at the renderer:
//
// documented — the long edge is a fact about the configured models: the strictest
// downscales to that size before reading, so rendering to it discards the pixels that
// model was going to discard anyway. On a deployment whose vision agents differ, a
// model with a larger long edge loses the difference. This is the same trade `imageLimitsHint` already
// recommends to a caller with an oversized IMAGE, applied by Iris to a page the
// caller never sized.
// assumed — the long edge is a guess, so rendering to it would be throwing away
// detail the model may well have read, which is the quiet damage this module exists
// to avoid. Only the hard ceiling is given up to, and only the pixels above it: they
// cannot be sent to the model at all under Iris's own rule (`dimensionReason`), so
// nothing that could have been read is lost. A page over the BYTE cap alone is
// therefore left to its rejection on an assumed basis — it is already inside the
// ceiling, so this returns null and the caller reports the refusal.
//
// Null for a page whose dimensions could not be read, too: the target is a comparison
// against the long edge this page HAS, and a page whose header would not parse has not
// told us. (pdftoppm writes a well-formed PNG, so this is a guard rather than a case.)
export function refitLongEdge(page: { width?: number; height?: number }, limits: ImageLimits): number | null {
if (page.width === undefined || page.height === undefined) return null;
const target = limits.basis === "documented" ? limits.max_long_edge_px : limits.max_dimension_px;
// A target at or above what the page already is would re-render the same picture and
// reject it twice.
return target < Math.max(page.width, page.height) ? target : null;
}

// Why a page is refused after Iris has already rendered it smaller, or null if the
// smaller render is fine.
//
// The same two limits and the same advice — the remedies in `rasterizedPageRejection`
// are still the caller's, and a page that is over at `longEdgePx` is over because of
// what is ON it. What this adds is the one thing that message would otherwise be wrong
// about: the dimensions it prints are the SECOND render's, not the one the DPI would
// have produced, and a caller comparing them against their own page would be measuring
// a picture they never asked for. Saying the retry happened also stops the obvious
// reply, which is to ask Iris to try a smaller size.
export function shrunkPageRejection(
pdfName: string,
pageNumber: number,
longEdgePx: number,
page: { bytes: number; width?: number; height?: number },
limits: ImageLimits,
): string | null {
const why = rasterizedPageRejection(pdfName, pageNumber, page, limits);
if (!why) return null;
// Whose number `longEdgePx` is, on the same terms as `dimensionReason`: on a
// documented basis `refitLongEdge` renders to what the model reads, and on an assumed
// one to the largest side Iris will send at all. Claiming the first where only the
// second is known would be putting a promise about the model in this sentence.
const size =
limits.basis === "documented"
? `${longEdgePx} px on the long edge, the size the vision model reads`
: `${longEdgePx} px on the long edge, the largest it will send`;
return `${why} Iris rendered this page again at ${size}, and it is still over.`;
}

// The long edge, in pixels, past which a rasterized page is bigger than the paper a
// document normally comes on. Letter at util/pdf.ts's 150 DPI is 1650 px and A4 is
// 1755; tabloid is 2550, and a drawing or a fold-out is larger still. Used only to
Expand Down
63 changes: 60 additions & 3 deletions src/routes/sessions.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,14 @@ import { runPipeline } from "../pipeline/orchestrator.ts";
import type { AuthedRequest } from "../auth/middleware.ts";
import { sendError } from "./errors.ts";
import { summarizeRun } from "../diagnostics.ts";
import { rasterizePdf, PdfTooLargeError, MAX_PDF_PAGES, type PageImage, type PdfLink } from "../util/pdf.ts";
import {
rasterizePdf,
rasterizePageToFit,
PdfTooLargeError,
MAX_PDF_PAGES,
type PageImage,
type PdfLink,
} from "../util/pdf.ts";
import { outputBasenameFromUploads, convertedHtmlFilename, safeStem, titledAs } from "../util/outputNames.ts";
import { captureFixtures } from "../pipeline/regression.ts";
import type { Fragment } from "../pipeline/fragment.ts";
Expand All @@ -29,7 +36,9 @@ import {
IMAGE_MEDIA_TYPES,
imageRejection,
rasterizedPageRejection,
refitLongEdge,
resolveImageLimits,
shrunkPageRejection,
} from "../providers/imageLimits.ts";
import { imageDimensions } from "../util/imageSize.ts";
import { readFields, tagPdf, taggedPdfCommand, tagTimeoutSeconds, TaggedPdfError } from "../util/taggedPdf.ts";
Expand Down Expand Up @@ -336,6 +345,10 @@ export function sessionsRouter(cfg: IrisConfig, store: Store): Router {
// carrying the link annotations on that page — the one part of a PDF that
// rasterizing destroys, so it travels alongside the image (see pipeline/links.ts).
const pages: PageImage[] = [];
// Pages that only fit after a second, smaller render. Collected rather than logged
// here because the session whose log this belongs in does not exist yet — the pages
// have to be measured before there is anything to record against (below).
const refits: { pdf: string; page: number; long_edge_px: number; from: string }[] = [];
try {
for (const f of files) {
if (PDF_EXT.test(f.originalname)) {
Expand All @@ -346,14 +359,45 @@ export function sessionsRouter(cfg: IrisConfig, store: Store): Router {
// what the model accepts — and the run would then die inside the first
// vision call, minutes in, which is the failure this route exists to catch.
for (const [i, p] of rendered.entries()) {
const page = p.page ?? i + 1;
const size = imageDimensions(p.buffer);
const why = rasterizedPageRejection(
f.originalname,
i + 1,
page,
{ bytes: p.buffer.length, width: size?.width, height: size?.height },
imageLimits,
);
if (why) throw new PageTooLargeError(why);
if (!why) continue;
// Iris chose these pixels, so before refusing the document it renders the
// page again at a size it can send (issue #485). Only this page, and only
// on the path that was about to fail: a document that converts today is
// rendered exactly as it was.
const target = refitLongEdge({ width: size?.width, height: size?.height }, imageLimits);
// A target means the page's dimensions parsed (`refitLongEdge` answers null
// otherwise), and `page` is set on everything `rasterizePdf` returns — but a
// rejection is the safe reading of either being absent, since re-rendering
// needs both the page to ask for and a size to ask for it at.
if (target === null || !size || p.page === undefined) throw new PageTooLargeError(why);
// A failed re-render gets the refusal it would have had, not pdftoppm's error.
const smaller = await rasterizePageToFit(f.buffer, p.page, target).catch(() => {
throw new PageTooLargeError(why);
});
const shrunk = imageDimensions(smaller);
const still = shrunkPageRejection(
f.originalname,
page,
target,
{ bytes: smaller.length, width: shrunk?.width, height: shrunk?.height },
imageLimits,
);
if (still) throw new PageTooLargeError(still);
refits.push({
pdf: f.originalname,
page,
long_edge_px: target,
from: `${size.width}x${size.height}`,
});
p.buffer = smaller;
}
pages.push(...rendered);
} else {
Expand Down Expand Up @@ -399,6 +443,19 @@ export function sessionsRouter(cfg: IrisConfig, store: Store): Router {
writeFileSync(paths.sessionLinks(sessionId), JSON.stringify(linksByOrder, null, 2));
}
keepSourcePdf(cfg, paths, sessionId, files);
// A page the model reads at lower resolution than the rest is a fact about this
// document's output, so the session says so rather than the deployment's stdout: the
// owner of a document whose fold-out reads worse than its letter pages can find out
// why from the log they already have (`GET /v1/sessions/{id}/logs`). Best-effort for
// the same reason `run_queued` is — a run must not fail to start over its own log.
if (refits.length) {
try {
const log = new RunLog(paths.sessionLog(sessionId));
for (const r of refits) log.event("page_refit", r);
} catch {
// ignore — observability must not block the run
}
}

const record = store.createSession({
session_id: sessionId,
Expand Down
57 changes: 56 additions & 1 deletion src/util/pdf.ts
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,10 @@ export interface PageImage {
// four pixels-worth of words with no target. Extracted separately and handed to
// the page agent as ground truth (see pipeline/links.ts).
links: PdfLink[];
// Which page of the PDF this image is, by the document's own numbering — what
// `rasterizePageToFit` needs to render it again. Absent for an uploaded image,
// which is not a page of anything.
page?: number;
}

// Thrown when a PDF exceeds the page cap, so the route can return a clean 400.
Expand Down Expand Up @@ -172,7 +176,7 @@ async function extractPdfLinks(pdfPath: string): Promise<Map<number, PdfLink[]>>
}

// Shards currently rendering, across every upload this process is serving. Read and
// written only by `rasterShards` and `rasterizePages`, and only between synchronous
// written only by `rasterShards`, `rasterizePages` and `rasterizePageToFit`, and only between synchronous
// statements — Node runs one of those at a time, so the reserve-then-spawn in
// `rasterizePages` cannot interleave with another document's and hand out the same
// cores twice.
Expand Down Expand Up @@ -344,8 +348,59 @@ export async function rasterizePdf(pdf: Buffer, originalName: string): Promise<P
name: `${base}-p${i + 1}.png`,
buffer: readFileSync(join(dir, f)),
links: links.get(pageNum(f)) ?? [],
page: pageNum(f),
}));
} finally {
rmSync(dir, { recursive: true, force: true });
}
}

// Render ONE page again, scaled so that neither side exceeds `longEdgePx`.
//
// Called only for a page whose render at `DPI` is over what the deployment can send
// (providers/imageLimits.ts): the alternative for that page is refusing the whole
// document, so fewer pixels beats no document, and which pixels may be given up is the
// limits module's decision rather than this one's (`refitLongEdge`). Nothing that
// converts today comes through here.
//
// `-scale-to` replaces `-r`: it fits the page inside a longEdgePx box preserving the
// aspect ratio, so what comes back is the same page rendered smaller — never cropped,
// and never a different page's ink. The page number is the PDF's own (`PageImage.page`),
// which is what `-f`/`-l` count in.
export async function rasterizePageToFit(pdf: Buffer, page: number, longEdgePx: number): Promise<Buffer> {
const dir = mkdtempSync(join(tmpdir(), "iris-pdf-fit-"));
try {
const pdfPath = join(dir, "in.pdf");
writeFileSync(pdfPath, pdf);
const out = join(dir, "pg");
// Reserved out of the host's render budget like any other shard (see
// `shardsRunning`): this is a pdftoppm process, and a count blind to it would hand
// the core it is using to the next document as well.
shardsRunning += 1;
try {
await execFileP("pdftoppm", [
"-png",
"-scale-to",
String(longEdgePx),
"-f",
String(page),
"-l",
String(page),
pdfPath,
out,
]);
} finally {
shardsRunning -= 1;
}
const pngs = readdirSync(dir).filter((f) => f.endsWith(".png"));
// One page in, one image out. Anything else means the range did not mean what this
// thinks it means, and silently returning the first file would hand the caller
// another page's ink under this page's name.
if (pngs.length !== 1) {
throw new Error(`re-rendering page ${page} produced ${pngs.length} images, expected 1`);
}
return readFileSync(join(dir, pngs[0]));
} finally {
rmSync(dir, { recursive: true, force: true });
}
}
Loading
Loading