Skip to content

Perceptual quality controller - #2

Merged
liminalism merged 27 commits into
mainfrom
perceptual-quality-controller
Aug 28, 2026
Merged

liminalism merged 27 commits into
mainfrom
perceptual-quality-controller

Conversation

@liminalism

Copy link
Copy Markdown
Owner

No description provided.

GitHub Uploader and others added 27 commits August 25, 2026 00:52
codegraph serve resolves the workspace from its own cwd; the explicit
--path ${workspaceFolder} argument was passed through unexpanded by some
launchers and broke indexing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
The outside advisor's answers to the one-shot brief (verified findings,
PR 0-5 program, gates). Registered via AKR source tooling; the root-level
working copy was byte-identical and is removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
…--quality seed

The one-shot program's policy-side PRs, per the registered 2026-08-24
advisor memo:

- PR 1: jpxl.quality-trace/2 groundwork — QualityWork counters (pixel
  plans, reconstructions, metric evaluations, entropy trainings,
  emissions) threaded through probe/price/rescue/reducer and folded
  across trials; prediction and transform_features blocks on
  QualityOutcome.
- PR 2: sweep_frame_perceptual — fresh production-policy pixel plans at
  pinned effective-scale rungs for oracle labelling; effective_scale /
  rung_for_effective_scale made pub.
- PR 4: TransformFeatureSummary — ten quantizer-independent DCT8
  features reduced per LF group in parallel through the shared candidate
  forward cache (deterministic across worker counts; in-search the
  prefill warms the cover's own cache and measured net-neutral wall).
- PR 5: the generated crossing model (qpv2-st-1, source + transform
  features, 130 image families) seeds the navigator — risk-adjusted
  candidate when confident, median when uncertain, legacy table only
  out of distribution — and its local slope prior steers the first
  correction under unchanged probe/price caps and canonical
  verification.

one-shot-controller joins the crate's default features after the
promotion A/B on 441 never-tuned holdout cells: zero floor violations
on either arm, byte geomean 0.997 (locked holdout) / 0.990 (painting
families), reconstructions -18% and wall -16% on the paintings, 12 MP
anchor wall -21%. Also closes the PR 7 reducer question's stale docs
and corrects the probe-cap accounting docs (observable maximum is
pixel_probes + 1 including the rescue probe).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
…quality-ladder

PR 0 of the one-shot program: an encode whose bounded search cannot
verify the requested score now fails at the public surface —
Error::TargetNotMet carrying a QualityMiss (kind, requested and best
verified scores, probe/price counts, metric version, trace), CLI exit 1
with no output file — instead of writing an under-target stream with
exit 0. Encoder::with_quality_fallback / --quality-fallback opt into
the two explicit alternatives: lossless (mathematically lossless
stream, status fallback_lossless) and best-effort (the previous
behavior, now explicit, under its true status). A refused encode still
appends its quality-trace record so failed searches remain calibration
input; every met-path stream is byte-identical.

Calibration surfaces for PR 1-4: the trace is now jpxl.quality-trace/2
(work, prediction and transform_features blocks; /1 stays readable by
the harness); new 'jpxl quality-ladder' subcommand (fresh
production-policy plans at pinned effective scales, canonically scored,
optionally exact-priced, JSONL); 'jpxl features --transform-summary'.
The CLI forwards the policy crate's one-shot-controller feature.

Decision record: jpegxl-rs.decision.quality-miss-fallback-semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
…ake, promotion A/B

The one-shot program's PR 2-3 calibration pipeline and the promotion
instrument:

- quality_oracle_labels.py: resumable ladder sweeps (sweep) and
  censoring-aware label extraction (labels) — coarsest-meeting-rung,
  log-loss interpolated crossing, local slope, priced neighbors;
  --time-budget-minutes gives a clean resumable stop for time-capped
  corpus runs.
- quality_predictor_v2.py: the pooled ln-loss crossing trainer —
  IRLS smoothed-pinball quantiles, robust standardization, ridge by
  held-out p90, family-grouped CV (LOFO or salted 10-fold), a blind
  never-tuned ext-holdout split excluded from tuning and from the
  shipped model, route simulation, and Rust codegen for the policy
  crate's generated predictor (run cargo fmt after regenerating).
- quality_corpus_extend.py: external-corpus intake — family grouping
  from filename stems, deterministic sha256 cal/dev/ext-holdout split,
  round-robin per-directory sampling, PPM conversion, a manifest under
  jpxl.codec-corpus-ext/1 (originals are never modified).
- one_shot_promotion_ab.py: interleaved A/B of two jpxl builds over
  manifest splits — floor violations, byte geomean, work counters,
  wall — the memo section 11 promotion screen.
- make-quality-guard-fixtures.py: family identity fields
  (family_id/variant_id/generator_family/source_capture_id), a guard
  that fails the build if any family crosses a split, and nine new
  fixtures (56 images / 38 families).
- codec_compare.py: reads quality-trace/2 (and still /1), passes
  --quality-fallback best-effort so RD sweeps grade under-target cells
  instead of dying on the new hard floor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
Ledger for the memo-driven session: the quality-miss fallback-semantics
decision; evidence for the falsification gate (honest FAIL on the
38-family corpus: LOFO median .22/p90 .81), the expanded-corpus gate
(blind painting holdout median .113/p90 .258, PASS) and the promotion
A/B (441 never-tuned cells, zero floor violations, PASS); the
pqc-one-shot-controller work record at rev 2 with the open items
(12 MP wall anchor 6.31x vs the 2.0x requirement, saturated/banding
corpus coverage, 108 unswept ext-polart images); the proposed JPEG
bitstream-recompression work record; and the regenerated views.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
Five gated phases from ISO/IEC 18181-2 s9.11 + Annex A and 18181-1: a
new clean-room jpxl-jpeg coefficient codec (self-gated byte-exact JPEG
round-trip), YCbCr 8x8-only VarDCT coefficient carriage with RAW
dequant matrices, the jbrd box + Annex A reconstructor, CLI/facade
surface, and black-box cjxl/djxl interop. Records the verified encoder
prerequisites (no do_YCbCr / jpeg_upsampling / RAW-quant writing today)
and the open questions (LF/DC bijectivity proof is the highest risk;
10918-1 source acquisition; OCR-suspect jbrd tables to check against
the original scan; Brotli dependency decision).

The work record jpegxl-rs.work.jpeg-bitstream-recompression now
carries the plan as origin source and the four phase gates as
acceptance checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
… on large frames

The 12 MP quality-mode wall was dominated by two serial per-probe
stages inside render_metric: the varblock reconstruction (a 12 MP
frame is a single LF group, so group-level parallelism gave it
nothing) and the transfer-curve linearization. Both are independent
per-element maps, so both now run banded over the EncodeExecutor with
byte-for-byte identical output at any worker count:

- varblocks: per-varblock compute (I.5.3/I.6/I.8/I.9 + J.4.3 sigma)
  moved into map_ordered chunks of 2048; each chunk's samples and
  sigma writes scatter serially into disjoint frame regions, and the
  dequant-matrix cache is warmed up front so the closure reads it
  through a shared borrow. Chunking bounds transient memory to a few
  MB; the locked 12 MP peak-memory anchor is unchanged (1,899 MB).
- linearization: into_linear_rgb_at_depth_with bands the per-sample
  LUT lookup; the large-frame scoring path in jpxl-perceptual uses it.

Verified byte-identical streams before/after on six image/target
cells, identical canonical scores, cross-thread determinism (threads
1/4/8, fast-debug and release), and green fmt/clippy/tests on the
touched crates. Release wall on the 12 MP anchor at 8 threads:
4.92 s -> 3.91 s (-21 %); matched-rate ratio 3.11x -> 2.46x on the
measuring host. The remaining per-probe residual is split between the
varblock scatter and the already-parallel metric.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
New observation pqc-large-frame-render-parallel-2026-08-25 (phase
attribution, exact wins, -21% release wall, ratio 3.11x -> 2.46x on
the measuring host, unchanged peak memory) and
pqc-usable-efforts-cost rev 6: the wall check remains the sole open
item, with the ranked residual directions — exact scatter overlap,
blur fusion (risky), and the score-changing reduced-pyramid
navigation that would need the Contract B promotion screen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
The jpxl-jpeg crate glob matched no tracked path (V-102) — the crate
does not exist yet; the plan document carries the intended layout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Tct9rZempR3FTEksR9t3VX
The per-chunk varblock results were scattered into the frame serially
after the parallel per-varblock compute. Each chunk's samples are now
bucketed by the 64-row bands its varblocks intersect (at most two per
varblock) and written by the executor's band workers; each sample
belongs to exactly one varblock and one band, so the placed pixels are
byte-for-byte the serial render at any worker count. Sigma writes are a
few hundredths of the sample volume and stay serial.

Measured honestly: wall-neutral at the 12 MP anchor (render_metric
2043/2168 ms before vs 2107/2057 ms after at 8 threads; 4 MP varblock
stage 45 ms at 4 threads in both trees) — the ~50-100 ms/probe scatter
estimate in the work note was high; the serial residual was ~30-45 ms
and inside run noise. Kept because it removes the last serial per-sample
stage, which grows with frame area (~150 ms/probe at 50 MP), at no
measured cost. Byte-identity verified against the pre-change binary on
six image/target cells and across threads 1/4/8 and AVX2 on/off; the
full 52-cell locked production-identity matrix passes (52/52 identical).

AKR-Change: chg-fff37cc35072bbe3
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Graph: sha256:55e79d09551a4c24e5d889fe846ff51f9fe0154a11423475bc7554fc025ffcfd
AKR-Tree: d43a520
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
Closes the one-shot memo's section-10 generation gap for the two
critical classes where worst-cell byte regressions concentrated:
saturated (4 -> 26 families) and gradient/banding-stress (8 -> 25).
Every new family is a distinct deterministic (seeded) content process —
hue wheels, angled saturated bands, near-clipping chroma ramps,
complementary checkerboards, saturated noise, sub-kilobyte-encodable
solid tiles, concentric rings, out-of-gamut clip plateaus, Voronoi;
shallow luma/chroma ramps, dithered ramps, radial sky glows, sunset
stops, bilinear fields, conic sweeps, diamond gradients — not one
recipe with 25 seeds. Corpus grows 56 -> 95 fixtures (calibration 43,
development 32, holdout 20); the split audit shows no family in two
splits, regeneration is sha256-idempotent, the manifest validates, and
all 13 previously locked holdout fixtures are byte-unchanged, so
existing evidence stays comparable. 16 generator unit tests pass.
Label sweeps over the new families are deliberately not run here.

AKR-Change: chg-c88603675047fe63
AKR-Work: jpegxl-rs.work.pqc-one-shot-controller
AKR-Graph: sha256:55e79d09551a4c24e5d889fe846ff51f9fe0154a11423475bc7554fc025ffcfd
AKR-Tree: 92aa67c
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
- README: verified against CONFORMANCE and the CLI; the quality-search
  section now describes the promoted one-shot seed (qpv2-st-1 pooled
  quantile regression behind the default `one-shot-controller` feature,
  with the calibrated-table fallback for out-of-envelope inputs) and the
  full-resolution canonical verification before every emission.
- JPXL/docs/LICENSING-AUDIT.md: MIT-only audit of all 12 workspace
  crates and the full dependency graph; no copyleft anywhere; the only
  non-MIT arms are permissive (moxcms/pxfm BSD-3/Apache-2.0 on the
  default path via the image adapter; BSD/CC0 metric crates behind
  off-by-default measurement features).
- JPXL/tools/bench_vs_libjxl.sh: dependency-light reproducible
  size/SSIMULACRA2/wall comparison (jpxl vs cjxl/djxl) with provenance
  headers (versions, sha256, flags), JSONL output, and clearly marked
  JPXL-only degradation when no runnable native oracle is present.
  Smoke-tested; native Linux cjxl/djxl (setup-oracles.sh) is still
  needed to exercise the comparison columns on this host.

AKR-Change: chg-3c1cf97d55c892ef
AKR-Work: jpegxl-rs.work.publish-readme-benchmark-mit
AKR-Graph: sha256:55e79d09551a4c24e5d889fe846ff51f9fe0154a11423475bc7554fc025ffcfd
AKR-Tree: cdc6873
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
New clean-room crate implementing ISO/IEC 10918-1 at the coefficient
level: markers, quant and Huffman tables, frame/scan headers, and
entropy-decoded DCT coefficients, re-emitted byte-for-byte. Baseline
(SOF0), extended sequential (SOF1), and progressive (SOF2, all four
procedures) are covered; restart intervals capture per-segment padding;
APPn/COM/unknown markers and post-EOI tail data are preserved verbatim;
arithmetic, hierarchical, lossless, and 12-bit inputs are refused with
a typed error. No workspace or third-party dependencies, no unsafe,
hand-rolled error enum, unit-bearing newtypes.

The one non-derivable wire fact is recorded rather than re-derived:
libjpeg splits long progressive AC end-of-band runs when its
correction-bit buffer fills, and that split is invisible in the
coefficients, so decode records the exact EOB-run lengths per scan and
encode replays them flush-on-match. A regression fixture with split
runs proves the recorded runs are load-bearing.

Verified over the full private Pol Art archive via the deterministic
roundtrip_sweep example: 5,217 of 5,218 JPEGs byte-identical, zero
mismatches (coverage: 3,166 baseline, 2,049 progressive, 2 extended
sequential, 4:4:4/4:2:0/4:2:2/4:4:0, grayscale, 684 restart-interval,
Exif/ICC/XMP, 41 with trailing data); the one remaining file is a
truncated zero-padded JPEG that libjpeg's djpeg also rejects, refused
here with a typed error. 34 crate tests pass (17 bit-exact roundtrips
on generated fixtures with provenance sidecars, 10 typed refusals,
2 EOB-run regressions, 4 unit, 1 doctest); workspace gates green.

AKR-Change: chg-2089f6782d8c12a3
AKR-Work: jpegxl-rs.work.jpeg-bitstream-recompression
AKR-Graph: sha256:6eb74db06baef609d12d358f1a48e5fe3a85a55f9db79ea04ef192b562d021f7
AKR-Tree: 6b253d1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
- pqc-usable-efforts-cost rev 7: memory-12mp (1.95 GiB peak) and
  production-identity (52/52 cells, threads 1/4, AVX2 on/off) re-verified
  on the current binary; the quiet-host wall observation re-baselines the
  anchors (12 MP 5.42x/4t, 4.25x/8t) and shows the 2.0x wall target is
  unreachable by search improvements alone - fixed costs are ~1.95x the
  rate path; wall-anchors stays open pending a target decision.
- pqc-one-shot-controller rev 3: the saturated/gradient corpus generation
  gap is closed (26/25 families); label sweeps remain.
- jpeg-bitstream-recompression active: Phase A evidence over the full
  5,218-file archive; the EOB-run record/replay discovery is noted for
  Phases B/C.
- publish-readme-benchmark-mit completed: all four checks evidenced
  (reproducible native-oracle comparison, MIT-only audit, README
  current-state, workspace gates).
- papercut: akr commit-msg hook rejects message-only amends after the
  change transaction closes.

AKR-Change: chg-186f2acbbc075c3f
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Evidence: @jpegxl-rs.evidence.bench-vs-libjxl-reproducible-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.jpeg-phase-a-roundtrip-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.mit-only-licensing-audit-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-memory-12mp-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-scatter-production-identity-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.readme-current-state-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.workspace-gates-2026-08-25/1
AKR-Graph: sha256:30546cc5534a77086da577a50522e7178dea87561649cc48bd29bbec56746c56
AKR-Tree: 3511116
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
Profile and reduce the production Fast and Balanced perceptual path's
full-frame render/metric allocation and rescue-probe cost. Work only on
usable efforts: do not spend measurement time on the feature-gated
Quality reference effort. Pure scorer, renderer, and lifetime changes
must preserve Fast/Balanced codestream bytes; any deliberate
search-policy change requires the standing Contract B screen.

AKR-Change: chg-9465cc41a44894e8
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Graph: sha256:30546cc5534a77086da577a50522e7178dea87561649cc48bd29bbec56746c56
AKR-Tree: ed9e0af
- jpegxl-rs.papercut.a-work-revision-that-cites-already-committed new -> verified

No AKR work record: ledger maintenance: papercut record and regenerated
views only

AKR-Change: chg-d73accb685d2f11e
AKR-Graph: sha256:ef95ff5110be9aeb247bd9805fd425acc01148ef91bb9ca4e8c9dd0c7efb3f8c
AKR-Tree: 7fceaad
…paths

Four value-preserving optimizations in the SSIMULACRA2 probe path, found
by phase attribution at the 12 MP anchor (per probe at 8 threads: render
195 ms, score 290 ms, of which scale-0 candidate blurs ~160 ms):

- The vertical blur pass writes each column strip's disjoint columns of
  the output in place (DisjointColumns), replacing per-strip buffers, a
  mutex collection, and a serial full-plane scatter copy. Per scale-0
  blur the vertical half dropped from ~10 ms to ~4.7 ms.
- The horizontal blur pass runs four rows at a time as independent
  recursion lanes over a lane-interleaved padded row, the row analogue
  of the vertical pass's column lanes, dispatched to an AVX2 build like
  the vertical pass. Per lane the arithmetic is exactly the scalar
  step's, in the same order.
- Redundant re-zeroing is dropped where every element is provably
  overwritten before it is read: the blur's horizontal temp (re-memset
  up to nine times 48 MB per 12 MP probe) and the mu2/s22/s12 and
  recomputed-reference scratch planes.
- from_srgb16 evaluates the transfer function through a 65536-entry
  table of the same expression instead of a powf per sample, and
  deinterleave preallocates its planes.

Exactness: 6/6 A/B cells byte-identical vs the pre-change binary
(threads 1/4/8, AVX2 on/off), full locked production-identity matrix
52/52 identical, 12 MP output bytes unchanged. Measured on the standing
interleaved recipe (load_1m 1.8-4.5): quality/rate 3.40x/2.76x at
4.3 MP (4t/8t) and 5.07x/3.50x at 12 MP, from 3.87x/3.22x and
5.42x/4.25x; 12 MP peak RSS 2,008,188 KiB from 2,043,056 (ceiling
2,097,152). Artifacts: .agent/scratch/pqc-metric-cuts-20260825 (kept).

AKR-Change: chg-24eecb5fd4bda387
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Graph: sha256:ef95ff5110be9aeb247bd9805fd425acc01148ef91bb9ca4e8c9dd0c7efb3f8c
AKR-Tree: a9240f5
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
Profile and reduce the production Fast and Balanced perceptual path's
full-frame render/metric allocation and rescue-probe cost. Work only on
usable efforts: do not spend measurement time on the feature-gated
Quality reference effort. Pure scorer, renderer, and lifetime changes
must preserve Fast/Balanced codestream bytes; any deliberate
search-policy change requires the standing Contract B screen.

- jpegxl-rs.evidence.pqc-metric-cuts-memory-12mp-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-metric-cuts-production-identity-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-metric-cuts-wall-anchors-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-metric-cuts-workspace-gates-2026-08-25 new -> verified

Verified by:
- Metric-cuts 12 MP memory: peak 2,008,188 KiB under the 2 GiB ceiling
- Metric-cuts production identity: 52/52 matrix and 6/6 pre-change A/B byte-identical
- Metric-cuts wall anchors: ratios cut ~9-19 percent, 2.0x target still unmet
- Metric-cuts workspace gates: fmt, clippy -D warnings, release tests all pass

AKR-Change: chg-ae2be006272a0327
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Evidence: @jpegxl-rs.evidence.pqc-metric-cuts-memory-12mp-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-metric-cuts-production-identity-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-metric-cuts-wall-anchors-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-metric-cuts-workspace-gates-2026-08-25/1
AKR-Graph: sha256:11b87518a2e357f90ab0354254cb1ca3fef23f1be56d7cd245257f95d636a6a5
AKR-Tree: 7036c46
- jpegxl-rs.papercut.knowledge-get-with-detail-canonical-on-jpegxl new -> verified

No AKR work record: papercut-only ledger change names no work record

AKR-Change: chg-cc21e841d39f4cd6
AKR-Graph: sha256:c5469183a640fdfb5dd28a9a6034632a0c2f9e378d5469f0dd6a6c1dd597d679
AKR-Tree: a9df958
…pth-quantized probe render

Guided by a real cycle profile (perf, enabled by the operator): after the
blur cuts, `__powf_fma` was 5.4% of cycles — a per-sample sRGB `powf`
encode in the render's colour pass whose result every probe immediately
quantized away, plus a per-sample transfer `powf` in the 16-bit prep —
and the two-step linearize pass another ~3%.

- PlanRenderer::render_linear_at_depth_with renders straight to the
  depth-quantized linear planes a probe scores: each linear sample's
  integer level is found by comparing against thresholds bisected over
  the f32 bit lattice from the actual linear_to_srgb -> scale -> round
  -> clamp composition, then mapped through the existing linear table.
  A guide table over the bit lattice pins each sample's level to a
  bucket-local range (two loads and at most a short ordered scan), and
  the classifier is cached on the renderer per bit depth so thresholds
  are bisected once per search, not per probe. The signalled sRGB
  encoding is never materialised on this path; render_with and
  RenderedFrame are unchanged for every other caller.
- The classifier's equivalence with encode-then-round-trip is pinned by
  fused_depth_levels_match_the_srgb_round_trip: a dense sweep, every
  threshold's bit-lattice neighbourhood, both transfer branch seams, and
  the non-finite specials (which take level 0 exactly as the round
  trip's finiteness guard does).
- PlanRenderEvaluator scores through the fused path on both the standard
  and the large-image branches.
- PreparedFrame::from_srgb16_with evaluates its transfer function
  through a 65536-entry table of the same per-value expression.

Exactness: 6/6 A/B cells byte-identical vs the pre-change 1fc3ca7
binary (threads 1/4/8, AVX2 on/off), locked production-identity matrix
52/52 identical, 12 MP bytes unchanged. Quiet-host 12 MP 8-thread
quality wall 2.39-2.53 s (2.62 s before the fusion); interleaved-recipe
ratios 3.06x/2.50x at 4.3 MP (4t/8t) and 4.65x at 12 MP/4t, from
3.40x/2.76x and 5.07x this morning (the 12 MP/8t schedule cell was
load-polluted; see the kept scratch README). Peak 12 MP RSS unchanged.
Artifacts: .agent/scratch/pqc-metric-cuts-20260825 (round 2).

AKR-Change: chg-950455e93e20d5ce
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Graph: sha256:c5469183a640fdfb5dd28a9a6034632a0c2f9e378d5469f0dd6a6c1dd597d679
AKR-Tree: c6a9fe7
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PGFGWnmbi5WRCwCLu7s6K
Profile and reduce the production Fast and Balanced perceptual path's
full-frame render/metric allocation and rescue-probe cost. Work only on
usable efforts: do not spend measurement time on the feature-gated
Quality reference effort. Pure scorer, renderer, and lifetime changes
must preserve Fast/Balanced codestream bytes; any deliberate
search-policy change requires the standing Contract B screen.

- jpegxl-rs.evidence.pqc-fused-render-memory-12mp-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-fused-render-production-identity-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-fused-render-wall-anchors-2026-08-25 new -> verified
- jpegxl-rs.evidence.pqc-fused-render-workspace-gates-2026-08-25 new -> verified

Verified by:
- Fused-render 12 MP memory: peak 2,008,408 KiB under the 2 GiB ceiling
- Fused-render production identity: 52/52 matrix and 6/6 pre-change A/B byte-identical
- Fused-render wall anchors: 4.3 MP at 2.50x (8t), 2.0x target still unmet
- Fused-render workspace gates: fmt, clippy -D warnings, release tests all pass

AKR-Change: chg-65a35814006d248f
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Evidence: @jpegxl-rs.evidence.pqc-fused-render-memory-12mp-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-fused-render-production-identity-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-fused-render-wall-anchors-2026-08-25/1
AKR-Evidence: @jpegxl-rs.evidence.pqc-fused-render-workspace-gates-2026-08-25/1
AKR-Graph: sha256:cd475cd59072daec056b7dd990d7a577cabe3bed751c45724d18b3e51111151d
AKR-Tree: 0a22526
Add the two metadata seams the archive pipeline needs:

- append_exif writes a Part 2 Exif box (LBox + "Exif" + u32
  tiff_header_offset + raw TIFF payload) after the codestream boxes;
  the facade exposes Encoder::with_exif, which validates the TIFF
  byte-order header and implies the container.
- ColourSpace {Srgb, LinearSrgb, DisplayP3, Rec2020} threads through
  EncodeOptions into the Annex E ColourEncoding bundle (Table E.1
  non-default sequence). Lossless Modular signals colour declaratively;
  lossy VarDCT rejects non-sRGB with a typed Unsupported error because
  the XYB transform and the metric are defined on sRGB input.
- Facade wrapping is unified in finish(); the rgb8/gray8 paths no
  longer widen u8 samples through an intermediate Vec<u16>.

Round-trip tests cover the box walk, all four colour spaces through the
decoder, the greyscale guard, and TIFF-header rejection.

- jpegxl-rs.work.pipeline-metadata-seam new -> proposed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsAR4R1kGCRk8spmhQZMC2
AKR-Change: chg-2985272a860c003a
AKR-Work: jpegxl-rs.work.pipeline-metadata-seam
AKR-Graph: sha256:e56b9a85583615d94e1fab8b121fb543a4e24563fe57c169a07ccbb4dcc093fd
AKR-Tree: 7985656
accumulate_band now dispatches to an avx2+fma multiversioned clone when
the CPU supports it. The lane loop computes the per-pixel SSIM' and
asymmetry terms in POOL_LANES groups, then folds them into the f64
MapSums in original pixel order, so the accumulated sums -- and the
encoded bytes -- are identical to the scalar walk (verified by test
across remainder lengths and by cmp on full encodes). ~5-8% off the
12 MP quality-path wall.

jpxl-plan-render now names the jpxl-core "simd" feature it depends on
instead of relying on feature unification.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsAR4R1kGCRk8spmhQZMC2
AKR-Change: chg-35a2d60333479cdd
AKR-Work: jpegxl-rs.work.pqc-usable-efforts-cost
AKR-Graph: sha256:e56b9a85583615d94e1fab8b121fb543a4e24563fe57c169a07ccbb4dcc093fd
AKR-Tree: 825f2bb
tools/build-release-final.sh runs the measured PGO cycle (instrumented
build, training encodes over the canonical test set across the quality,
rate and lossless paths plus a decode, llvm-profdata merge, profile-use
rebuild) on the new release-final profile. Measured today the combined
fat-LTO+PGO gain over `release` is 2-8%; the ledger's 36-53% PGO
evidence predates the fused-render and Opt-F..P rounds and needs
re-verification. `release` (thin LTO) stays the benchmark profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsAR4R1kGCRk8spmhQZMC2
AKR-Change: chg-105dcc145ce18786
AKR-Graph: sha256:e56b9a85583615d94e1fab8b121fb543a4e24563fe57c169a07ccbb4dcc093fd
AKR-Tree: 58d9bd8
The metadata seam is verified end to end at this commit: an appended
Exif box round-trips through the decoder's box walk with the codestream
untouched; with_exif rejects payloads

- jpegxl-rs.evidence.avx2-pooling-bit-identity-2026-08-26 new -> verified
- jpegxl-rs.evidence.pipeline-metadata-seam-2026-08-26 new -> verified
- jpegxl-rs.observation.pgo-gain-remeasured-2026-08-26 new -> verified

Verified by:
- jpegxl-rs.evidence.avx2-pooling-bit-identity-2026-08-26
- jpegxl-rs.evidence.pipeline-metadata-seam-2026-08-26

AKR-Change: chg-438b67ff073613c8
AKR-Work: jpegxl-rs.evidence.pipeline-metadata-seam-2026-08-26
AKR-Work: jpegxl-rs.evidence.avx2-pooling-bit-identity-2026-08-26
AKR-Work: jpegxl-rs.observation.pgo-gain-remeasured-2026-08-26
AKR-Evidence: @jpegxl-rs.evidence.avx2-pooling-bit-identity-2026-08-26/1
AKR-Evidence: @jpegxl-rs.evidence.pipeline-metadata-seam-2026-08-26/1
AKR-Graph: sha256:2e53454b4abc5a2e7803995f3fb00745b536d02f33e33c9607223cdf0dd3935b
AKR-Tree: 3f08f2f
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QsAR4R1kGCRk8spmhQZMC2
Authoritative plan for jpegxl-rs.track.encoder-optimization: execute
Opt-F through Opt-P in after-order, supported by the static advisor
assessment and the optimization decisions and gates.

AKR-Change: chg-d328f28ddbc84320
AKR-Work: jpegxl-rs.work.encoder-optimization-plan
AKR-Graph: sha256:2e53454b4abc5a2e7803995f3fb00745b536d02f33e33c9607223cdf0dd3935b
AKR-Tree: d7a6de0
@liminalism
liminalism merged commit c849f21 into main Aug 28, 2026
0 of 2 checks passed
@liminalism
liminalism deleted the perceptual-quality-controller branch August 28, 2026 09:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant