Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,8 @@

## Unreleased

- Add versioned retained-owner and finalization receipts without changing search policy or the checksum-bound observability-v1/v2 shapes. Shared pass telemetry now accounts for the bounded observability ledger through a non-recursive canonical UTF-8 projection and preserves its known high-water state across resume; older checkpoints remain readable and explicitly report incomplete owner history. Checkpoint saves report logical graph bytes, streamed artifact/chunk bytes, compressed output, configured compression capacity, durable payload/manifest bytes, and whether encode/compress work was applied or an existing stable ID was reused. Report saves similarly expose canonical identity-string, artifact-serialization, conservative potential string, and durable bytes outside the content-derived report. Configured finalization headroom is labeled unallocated capacity and remains excluded from retained/process totals. Stable-ID algorithms are unchanged: deterministic owner state participates in newly emitted checkpoint content, while post-write receipts never feed back into the IDs or payloads they measure. These fields are observational/accounting inputs only: they do not emit the reserved checkpoint/epoch/pressure reasons, change frontier order, activate a budget or stopping policy, add checkpoint v2/readback ownership, or publish a release; #156 and #216 remain open.
- Extend the partial owner ledger across checkpoint reopen and report enrichment/finalization without adding another payload read, decode, parse, or stable-ID traversal. Successful checkpoint opens and resumes expose bounded stage-by-stage receipts, while typed corrupt/unsupported/resource-limit failures retain the work reached before failure. The opt-in accounted report builder measures the source exploration graph, finding-identity strings already materialized by enrichment, and the returned report; report saves separately distinguish created envelope graphs from reuse-parsed artifact graphs. These compact-JSON serialized-view proxies and their conservative potential sums are not exclusive retained-owner bytes or observed heap peaks. Runtime receipts remain outside saved report/checkpoint payloads and IDs, and bounded JSON streaming still avoids the monolithic report graph. This adds no checkpoint v2/framed readback, allocation, eviction, compaction, stopping policy, benchmark restart, or release; #156 and #216 remain open.
- Add versioned retained-owner and finalization receipts without changing search policy or the checksum-bound observability-v1/v2 shapes. Shared pass telemetry now accounts for the bounded observability ledger through a non-recursive canonical UTF-8 projection and preserves its known high-water state across resume; older checkpoints remain readable and explicitly report incomplete owner history. Checkpoint saves report logical graph bytes, streamed artifact/chunk bytes, compressed output, configured compression capacity, durable payload/manifest bytes, and whether encode/compress work was applied or an existing stable ID was reused. Report saves similarly expose canonical identity-string, artifact-serialization, conservative potential string, and durable bytes outside the content-derived report. Configured finalization headroom is labeled unallocated capacity and remains excluded from retained/process totals. Stable-ID algorithms are unchanged: deterministic owner state participates in newly emitted checkpoint content, while post-write receipts never feed back into the IDs or payloads they measure. These fields are observational/accounting inputs only: they do not emit the reserved checkpoint/epoch/pressure reasons, change frontier order, activate a budget or stopping policy, add checkpoint v2, or publish a release; #156 and #216 remain open.
- Extend shared-search observability to nested schema v2 with monotonic sample sequences, canonical coalesced reason vectors, an exact 15-bit trigger mask, a fixed eight-count boundary-local trigger-yield tuple, category/frontier/cadence/termination triggers, and category-specific first/last discovery, current/longest dry-state, and identities-per-million-transition facts. The redundant vector/mask/tuple preserves exact retained-boundary triggers across compaction while aggregate interval deltas are rebuilt: adjacent complete-history samples require the tuple to equal the semantic interval delta, while genuine compaction gaps or incomplete v1 history allow only a componentwise-bounded tuple. Observability-v1 migration synthesizes masks only from explicit historical boundaries and uses a zero tuple to mean no v1 trigger was recorded, not that an interval was event-free. Sequence validation applies monotonic and feasible record/state/cadence bounds without claiming exact replay of discarded compaction history. Emitted v2 checkpoints add a domain-separated, key-order-independent canonical SHA-256 self-check binding every fixed known ledger field; it detects stale or accidental mutation but is not authentication because a writer can recompute it, and persisted payload/manifest digests remain the external artifact-corruption boundary. V1 and pre-ledger inputs gain the nested checksum on their first emitted v2 checkpoint. Live process sampling and progress output remain bounded to cadence and final termination while the compacted deterministic ledger retains event boundaries. Checkpoint construction remains side-effect-free and deterministic: the persisted ledger participates in stable checkpoint identity, while observability-v1 checkpoints migrate explicitly with incomplete event history and nullable facts rather than invented history. This remains an independent partial #216 slice: `checkpoint`, `epoch`, and generic `pressure` reasons are reserved but not emitted; retained-owner accounting remains incomplete; no resource, allocation, or stopping policy is activated; #216 remains open; and the earlier V1 overhead study is not an exact-head overhead claim for this V2 implementation.
- Add the initial bounded fixed-interval/termination `ResourceSampleV1` ledger for shared search with deterministic logical retention, category-specific interval yield, early-versus-post-first summaries, fail-closed exact-resume persistence, strict compact machine projection, privacy-safe live heap/RSS observations, and explicit monotonic run-wide positions across additive goal passes. The nested-v2 extension above adds category and frontier event boundaries; checkpoint, epoch, and generic pressure boundaries remain later work.
- Bound schema-v1 checkpoint readback by stored and decompressed bytes, classify reopen failures as corrupt, unsupported, or resource-limited, and add canonically self-bound private manifests so listing and retention use bounded metadata I/O without opening new frontiers. Full payload digests remain an open/resume boundary; no-clobber same-ID publication, payload-first crash recovery, and pair-inclusive quotas preserve existing v1 JSON/gzip reads, stable IDs, and exact resume pending framed schema v2.
Expand Down
4 changes: 3 additions & 1 deletion docs/evidence-ndjson.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,9 @@ Current events use `schemaVersion: 1`:
- `ending` records a stable event ID, global `elapsedMs`, pass-local `firstDiscoveredAtState`, numeric `choiceIndices`, and optional `foundBy` pass. It omits choice prose, final text, and variables.
- `runtime_error` records the same replay coordinates plus the error message and optional source location.
- `benchmark_signal` is reserved for InkBench's oracle-neutral fixtures. A root `INKBENCH_SIGNAL_MODE` tag suppresses ordinary finding records, and an exact `INKBENCH_SIGNAL:<non-negative integer>` tag emits only that numeric signal, its numeric replay path, and timing. No story prose, variables, or arbitrary tags are copied.
- `run_end` is the authoritative bounded summary. It records compile status, states explored, finding counts, limits, execution mode, truncation causes, resource envelopes/deadlines, and emitted-evidence counts. Its versioned `resources.ownerAccounting` labels configured finalization memory/time reserve as unallocated capacity, carries at most eight deterministic numeric shared-pass retained-owner summaries, and includes a post-write checkpoint receipt when checkpoint persistence occurred. Report-save accounting is carried by the ordinary JSON `artifact` reference because `--json-stream` and `--save-report` are mutually exclusive. It does not contain the full ending, pass, schedule, or discovery-curve arrays.
- `run_end` is the authoritative bounded summary. It records compile status, states explored, finding counts, limits, execution mode, truncation causes, resource envelopes/deadlines, and emitted-evidence counts. Its versioned `resources.ownerAccounting` labels configured finalization memory/time reserve as unallocated capacity, carries at most eight deterministic numeric shared-pass retained-owner summaries, includes checkpoint-reopen accounting on a resumed run, and includes a post-write checkpoint receipt when checkpoint persistence occurred. The bounded stream deliberately has no report-enrichment or report-finalization receipt because it neither constructs the monolithic report graph nor permits `--save-report`. It does not contain the full ending, pass, schedule, or discovery-curve arrays.

Successful ordinary `--json` output adds the same bounded `resources.ownerAccounting` wrapper only after the canonical report has been built. It may include checkpoint-read, report-enrichment, checkpoint-commit, and report-finalization receipts as those lifecycle stages apply. The wrapper and the `artifact` reference remain outside the report passed to the artifact writer, so neither the saved report payload nor its stable ID contains these receipts. The report-enrichment values are compact-JSON serialized-view proxies and a conservative potential sum; shared backing can appear in both the source and returned projections, so they are not exclusive retained-owner bytes or an observed heap peak.

Events are flushed as newline-delimited records during the run. A consumer may retain and replay complete finding lines even if an outer process guard later interrupts the CLI. Absence of `run_end` means the run was interrupted; it must not be relabeled as clean completion.

Expand Down
12 changes: 10 additions & 2 deletions docs/local-artifacts.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,19 @@ Inkcheck can persist a completed CLI report without uploading story source:
inkcheck story.ink --save-report --json
```

Saving is explicit. Without `--save-report`, Inkcheck creates no report artifact and its existing stdout contract is unchanged. With the flag, the JSON report adds an `artifact` reference and stderr confirms the same stable ID. The reference also carries a schema-v1 finalization `accounting` receipt: canonical report-ID input bytes, whether artifact serialization was applied, pretty-printed artifact UTF-8 bytes, reuse readback and validation-identity bytes, a conservative total for those strings, zero current string ownership after return, and durable file bytes. The report is stored under `.inkcheck/reports/report-<hash>.json` using a same-directory temporary file and atomic rename.
Saving is explicit. Without `--save-report`, Inkcheck creates no report artifact, emits no artifact reference, and prints no saved-artifact confirmation. Successful monolithic `--json` output may still add the numeric `resources.ownerAccounting.reportEnrichment` lifecycle receipt outside the report payload; bounded `--json-stream` deliberately does not build that full enriched graph. With `--save-report`, the JSON report additionally carries an `artifact` reference and stderr confirms the same stable ID. The reference carries a schema-v1 finalization `accounting` receipt. It reports canonical report-ID input bytes, whether artifact serialization was applied, pretty-printed artifact UTF-8 bytes, reuse readback and validation-identity bytes, a conservative total for those strings, zero current string ownership after return, durable file bytes, and separately discriminated graph-materialization proxies. The report is stored under `.inkcheck/reports/report-<hash>.json` using a same-directory temporary file and atomic rename.

The stable ID is derived from canonical report content plus its project-relative entrypoint binding. Repeating the same deterministic check on the same entrypoint reuses the same artifact instead of creating timestamp duplicates; identical reports from two different files cannot alias each other. Each versioned envelope records its creation time, Inkcheck and report schema versions, project-relative entrypoint, source fingerprint, effective configuration, and complete report.

The accounting receipt is deliberately outside the saved report and its content-derived ID. A reuse materializes the caller's schema-v1 canonical identity, reads the existing artifact into one raw UTF-8 string, and materializes a second canonical identity while validating the parsed report. It does not create a new pretty-printed artifact string: `serialization.status` is `not_applied` while `readback.status` is `applied`. `readback.rawArtifactString` reports that one string's count and logical UTF-8 bytes, and `readback.validationIdentity` reports the second canonical identity's count and bytes. A created artifact has `readback.status: "not_applied"` and zero-valued readback owners. The receipt derives readback bytes from the same validation read rather than reading or serializing the artifact again. Its conservative string total includes both identities and the raw artifact string; it does not claim an observed V8 heap peak or prove when garbage collection reclaimed an earlier string.
The accounting receipt is deliberately outside the saved report and its content-derived ID. Full report lifecycles may attach `reportEnrichment.status: "applied"`, containing numeric schema-v1 accounting for the source `ExploreResult` compact-JSON view, stable finding-ID input strings materialized during enrichment, the returned report compact-JSON view, their conservative potential sum, and the returned view still owned by the caller. Legacy callers that did not opt into the accounted builder receive the explicit `{ status: "not_applied", reason: "unavailable" }` variant; absence is not reported as a measured zero. The saver projects the numeric whitelist rather than copying arbitrary caller keys, so accounting does not duplicate authored prose, witness text, variables, or finding messages.

`graphMaterialization.createdEnvelope` accounts the compact-JSON view of the artifact envelope constructed by an accounted create. On an accounted reuse, `graphMaterialization.reuseParsedArtifact` instead accounts the artifact graph returned by the existing validation parse. Exactly one graph is `applied`; the other is `not_applied`. A legacy direct save without an enrichment receipt performs no new graph walk on either create or reuse and reports the outcome graph and `conservativePotential` as `not_applied`/`unavailable` rather than inventing zero bytes or changing pre-accounting behavior. Applied conservative potential stays separate from string materialization totals, and current saver-owned graph bytes are zero after the API returns. These deterministic serialized-view proxies are not observed heap/RSS peaks, exclusive retained-owner bytes, or proof of reclamation.

A reuse materializes the caller's schema-v1 canonical identity, reads the existing artifact into one raw UTF-8 string, and materializes a second canonical identity while validating the parsed report. It does not create a new pretty-printed artifact string: `serialization.status` is `not_applied` while `readback.status` is `applied`. `readback.rawArtifactString` reports that one string's count and logical UTF-8 bytes, and `readback.validationIdentity` reports the second canonical identity's count and bytes. A created artifact has `readback.status: "not_applied"` and zero-valued readback string owners. When reuse graph accounting is applied, its proxy is derived from that same read and parse rather than reading, parsing, or serializing the artifact again. Its conservative string total includes both identities and the raw artifact string; it does not include either graph proxy.

Graph byte counting streams JSON tokens without materializing a whole compact artifact string and accepts Inkcheck's schema-owned plain JSON containers only. It fails closed on cycles, `BigInt`, boxed/exotic or proxy containers, accessors, and inherited or own `toJSON` hooks rather than invoking active user code. Shared subgraphs are serialized once per reference, just as JSON would serialize them, so source-plus-returned potential deliberately double-counts shared DAG backing under a conservative no-sharing assumption. Do not add these proxies to shared-search retained-current or peak totals: charge-once backing attribution remains partial under issue #216.

MCP search-session persistence deliberately remains on the legacy unaccounted builder/saver path in this slice. That keeps enrichment and artifact graph walks out of memory-guard sampling before checkpoint suppression and `bindingLimit` decisions; session reports, IDs, and policy behavior therefore remain unchanged. The exposed accounted report lifecycle is opt-in, and its receipt remains outside the saved payload.

## Reopening safely

Expand Down
Loading