diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml new file mode 100644 index 0000000..82f0715 --- /dev/null +++ b/.github/workflows/ci.yml @@ -0,0 +1,21 @@ +name: CI + +on: + pull_request: + push: + branches: + - main + +permissions: + contents: read + +jobs: + test: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-node@v4 + with: + node-version: 20 + - run: npm test + - run: npm run conformance diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 71d3ea3..d972838 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -1,22 +1,45 @@ # Architecture ```text -raw operational output - ↓ -policy / parser / reducer - ↓ -bounded decision evidence - ↓ -model +caller-owned raw stdout/stderr + | + v + strict TAP-subset parser + | + v + evidence classes + verdict + | + v + deterministic ceiling policy + | | + v v + sufficient NEEDS_RAW_EVIDENCE + packet packet/control error ``` -## Ports +## Modules -- Input adapter: translates a host’s observable state into the generic contract. -- Core: deterministic policy/mechanism under test. -- Evidence store: immutable or append-only artifacts where required. -- Output adapter: returns a bounded receipt to the host. +- `src/reducer.js`: validation, byte framing, parsing, classification, policy, + hashing, measurement, and canonical serialization. +- `bin/context-firewall.js`: stdin/file CLI and deterministic conformance entry. +- `fixtures/corpus.js`: synthetic public-safe fixture definitions and expected + decisions. +- `tests/reducer.test.js`: success, failure, provenance, ambiguity, determinism, + measurement, and ceiling gates. -## Independence +The core has no host adapter or external package dependency. Supporting Opsle +protocols can consume the JSON fields without importing this package. -Host-specific adapters are optional and removable. Disabling the project should return the host to its prior behavior. No core module may import Taslos Tasks internals. +## Evidence ownership + +The caller owns storage and addressability of raw bytes. The reducer binds those +bytes to a digest and records the caller's reference. It neither stores nor +destroys the source. A later decision record can bind the reduced semantic hash, +input hash, configuration identity, and raw reference to state exactly what the +agent saw. + +## Deterministic boundary + +Caller-supplied duration is semantic input. Reducer execution latency is an +external measurement and must not enter the hashed packet. There is no clock, +randomness, network, locale sorting, or ambient host inspection in the core. diff --git a/BENCHMARK.md b/BENCHMARK.md index bfc13d9..708f715 100644 --- a/BENCHMARK.md +++ b/BENCHMARK.md @@ -33,3 +33,15 @@ Record model, provider, model version, reasoning effort, tool versions, fixture, ## Result policy Retain positive, negative, null, and failed experiments. Update maturity only when the actual stated hypothesis has reproducible evidence. + +## Prototype conformance is not EXP-001 + +`npm run conformance` executes synthetic parser/reducer boundary fixtures. Its +byte measurements demonstrate implementation behavior only. They are not model +correctness, a baseline/arm comparison, a frozen experimental dataset, or a +provider run, and must not be registered as EXP-001 results. + +Before EXP-001 can run, publish immutable task fixtures and hashes, a correctness +oracle, a harness with raw/reduced arms, exact model/provider configuration, and +randomized or blinded task allocation. Correctness must gate every efficiency +comparison. diff --git a/CHANGELOG.md b/CHANGELOG.md index c5f05f8..c9318aa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,15 @@ # Changelog +## 0.2.0 - 2026-08-25 + +- Added a dependency-free deterministic TAP-subset reducer and CLI. +- Added canonical evidence/provenance receipts, exact payload measurements, + source/configuration hashes, and explicit raw-evidence escalation. +- Added deterministic payload ceilings that fail closed when critical evidence + cannot fit. +- Added 30 synthetic conformance fixtures and 33 automated boundary tests. +- Promoted implementation maturity only to PROTOTYPED; EXP-001 remains PLANNED. + ## 0.1.0 - 2026-08-25 - Published initial theory, falsifiable specification, benchmark plan, architecture, and provenance. diff --git a/README.md b/README.md index 5ca89af..7fe9a48 100644 --- a/README.md +++ b/README.md @@ -2,66 +2,208 @@ > Experimental Opsle research. Claims are hypotheses until evidence supports them. -## Problem +## Thesis + +Operational output should not become model-visible merely because a tool emitted +it. Deterministic software can suppress repetitive success output while retaining +decision evidence, provenance, measurements, and a safe path back to raw bytes. + +The broader research question is: **How much context can an AI coding agent +safely not see?** This repository does not yet answer it. + +## Prototype scope + +Version 0.2.0 is a dependency-free Node.js reference reducer for a documented +flat TAP-compatible test-output subset. It: + +- reads caller-supplied stdout and stderr bytes plus process metadata; +- derives test verdict and pass, fail, and skip counts; +- retains every recognized failure header and failure-region line; +- retains strict fatal, timeout, and abnormal-warning markers; +- summarizes successful tests, structure, duration source lines, and explicit + informational notes; +- retains ambiguous or malformed evidence and requires raw-evidence escalation; +- emits canonical JSON with source/configuration hashes and payload measurements; +- applies deterministic payload ceilings without silently truncating critical + evidence. + +It does not parse arbitrary logs, Git output, compiler output, database output, +HTTP traces, or every test runner. It has no model, provider, network, database, +daemon, worker, UI, or product integration. It is not a production security +boundary. + +## Requirements + +- Node.js 20 or newer +- no package installation or network access + +## Input + +The CLI accepts one JSON object from stdin or `--input`. The input protocol is +`opsle.context-firewall.test-run-input/v1`. + +```json +{ + "protocol_version": "opsle.context-firewall.test-run-input/v1", + "operation_id": "op-public-example", + "source": { + "id": "synthetic-suite", + "run_id": "run-public-example", + "raw_evidence_ref": "artifact://public/run.tap" + }, + "process": { + "exit_code": 1, + "duration_ms": 12.5, + "interrupted": false + }, + "streams": [ + { + "name": "stdout", + "encoding": "utf8", + "data": "TAP version 13\nnot ok 1 - adds\n message: expected 2\n1..1\n# tests 1\n# pass 0\n# fail 1\n# skipped 0\n" + } + ] +} +``` + +Each stream is `stdout` or `stderr`; each may appear at most once. `encoding` is +`utf8` or `base64`. Base64 permits exact non-UTF-8 source bytes to be hashed and +reported without pretending they were classifiable text. + +Run the example: + +```bash +node ./bin/context-firewall.js reduce \ + --input examples/test-run-input.json +``` + +Read from stdin with a 4,096-byte ceiling: + +```bash +node ./bin/context-firewall.js reduce \ + --max-bytes 4096 +``` + +## Output packet + +The output protocol is +`opsle.context-firewall.evidence-packet/v1`. Canonical JSON contains: + +- `decision_evidence`: `passed`, `failed`, or `indeterminate` status; a + `SUFFICIENT` or `NEEDS_RAW_EVIDENCE` disposition; reason codes; process state; + aggregates; failures; fatal errors; timeouts; warnings; and unclassified data; +- `receipt.source`: caller-supplied source/run identity and per-stream byte + counts; +- `receipt.reducer` and `receipt.configuration`: exact reducer, policy, ceiling, + and configuration hash; +- `receipt.input_hash`: SHA-256 over framed stdout/stderr bytes; +- `receipt.semantic_payload_hash`: SHA-256 over canonical decision evidence; +- `receipt.measurements`: exact raw/reduced bytes and original, retained, and + suppressed event counts; +- `receipt.retained` and `receipt.suppressed`: explicit evidence taxonomy and + counts; +- `receipt.raw_evidence`: caller reference, escalation state, and the distinction + between context suppression and destruction/unavailability. + +The reducer never deletes raw evidence. A supplied raw reference is recorded as +`CALLER_REFERENCE_SUPPLIED`; the reducer does not claim it verified the external +artifact. Without a reference, preservation is `PRESERVATION_UNCONFIRMED` and +the packet requires escalation. + +`reduced_bytes` is the exact byte length of canonical stdout, including its final +newline. Raw and reduced bytes are sufficient for a trajectory consumer to +calculate visible fraction and reduction ratio without parsing human logs. +Runtime latency is intentionally absent from the hashed packet because it is +nondeterministic; callers may measure it outside the packet. + +## Deterministic retention policy + +Classification is strict and case-sensitive after ANSI is removed for parsing. +Original retained text still includes ANSI bytes. + +Priority is: + +1. verdict and process exit/interruption state; +2. fatal process/runner and timeout evidence; +3. every failed test identity, header, and failure-region line; +4. aggregate counts and supplied duration; +5. strict `WARNING:` or `WARN:` markers; +6. unclassified evidence; +7. structure, explicit `# note:` lines, and repetitive successful tests. + +Recognized failure regions are never partially truncated into a supposedly +sufficient packet. Under a ceiling, the reducer first replaces lower-priority +warning or unclassified text with hashes/locations and requires raw evidence. If +critical evidence still cannot fit, it emits a compact +`NEEDS_RAW_EVIDENCE` packet with failure identities and omits details only while +explicitly declaring the packet insufficient. If even that safe packet cannot +fit, stdout remains empty and the CLI emits a machine-readable +`PAYLOAD_CEILING_TOO_SMALL` error on stderr with exit code 2. + +The ceiling applies to model-visible stdout. Error-channel bytes are operational +control output and are not presented as a valid reduced packet. + +## Strict TAP subset and limits + +The parser recognizes: + +- `ok` and `not ok` records, optional numeric indexes, and `# SKIP`; +- indented or adjacent lines following `not ok` until the next test, plan, or + aggregate as one failure region; +- `# tests`, `# pass`, `# fail`, and `# skipped` aggregates; +- `TAP version`, plans, subtest headers, YAML delimiters, `# duration_ms`, and + `# note:` as structural/informational lines; +- strict `FATAL:`, `RUNNER CRASH:`, `UNCAUGHT:`, `TIMEOUT:`, `WARNING:`, and + `WARN:` markers outside failure regions. + +Unknown lines, TODO semantics, malformed UTF-8, contradictory aggregates, +interruption, missing exit state, missing source/operation identity, and +unexplained nonzero exits remain explicit and require raw evidence. Words such as +PASS, FAIL, error, and warning are not classified by substring. + +This flat subset does not establish full TAP conformance or support arbitrary +nested runner dialects. Small inputs commonly expand because the receipt has +fixed provenance cost; reduction is expected only for sufficiently repetitive +source payloads. + +## Verification + +Run all tests: + +```bash +npm test +``` + +Run the deterministic synthetic corpus: + +```bash +npm run conformance +``` + +The conformance result is canonical JSON over 30 public-safe fixtures spanning +normal success, normal failure, multi-failure retention, process problems, byte +and text edge cases, payload boundaries, and provenance gaps. It reports raw and +reduced sizes, verdict preservation, escalation, and PASS/FAIL. + +These are implementation fixtures, not model experiments or benchmark results. +They do not show that reduced context preserves agent correctness. + +## EXP-001 + +This prototype resolves one prerequisite for planned EXP-001: an executable, +deterministic reducer with synthetic conformance evidence. EXP-001 remains +PLANNED. Frozen experimental task fixtures, a correctness oracle, an experiment +harness, exact model/provider configurations, and randomized/blinded allocation +remain separate work. + +## Maturity and limitations + +**PROTOTYPED** under the canonical Opsle lifecycle. The implementation and its +boundary tests establish a runnable mechanism, not comparative correctness, +safe-frontier evidence, benchmark readiness, provider evidence, or replication. -Tools often send models large volumes of successful-test noise, shell chatter, and repeated state that does not affect the next decision. - -## Hypothesis - -Policy-driven reduction can remove substantial operational payload without lowering correctness when escalation and provenance remain available. - -## Mechanism - -Apply allow rules, suppression, aggregation, escalation, deterministic reducers, and payload ceilings between raw tool output and the model; retain links to auditable raw artifacts. - -## Why it matters - -The Opsle thesis asks: **What if we stopped using intelligence for work that doesn’t require intelligence?** This project isolates one candidate boundary so it can be falsified and measured independently. - -## Non-goals - -Hiding failures, lossy summarization without provenance, or assuming every tool can use the same fail-open/fail-closed policy. - -## Current maturity - -**THEORY** under the [Opsle maturity model](https://github.com/opsle/research/blob/main/MATURITY.md). - -## Existing evidence - -Bounded context construction and redaction patterns demonstrate feasibility. The safe reduction frontier is not established. - -## Evidence still missing - -Task-stratified correctness curves, reducer conformance suites, escalation policy, and adversarial omission testing. - -## Benchmark strategy - -Correctness gates every comparison. Planned measures: - -- correctness -- input bytes/tokens -- suppression ratio -- escalation rate -- missing-evidence defects -- reducer latency - -See [BENCHMARK.md](BENCHMARK.md) for experiment rules. No benchmark numbers are claimed. - -## Relationship to other Opsle research - -This project is part of [Opsle Research](https://github.com/opsle/research). Opsle Tasks is the future public name of the integrated reference system from which several ideas emerged. Its active development migration to the Opsle organization is intentionally deferred. - -## Relationship to future Opsle Tasks - -Future Opsle Tasks may consume this project through an adapter only after evidence supports integration. The active predecessor, Taslos Tasks, remains unchanged and has no dependency on this repository. - -## Installation status - -No installable production package is justified yet. The repository is theory/specification-first. - -## Known limitations - -Task-stratified correctness curves, reducer conformance suites, escalation policy, and adversarial omission testing. +See [THEORY.md](THEORY.md), [SPEC.md](SPEC.md), and +[BENCHMARK.md](BENCHMARK.md). ## License diff --git a/SPEC.md b/SPEC.md index 5aaa95e..a590f4d 100644 --- a/SPEC.md +++ b/SPEC.md @@ -1,38 +1,122 @@ -# Specification +# Context Firewall test-output reducer specification -Status: experimental theory contract. +Status: experimental prototype contract. + +Version: `opsle.context-firewall.evidence-packet/v1`. ## Compatibility boundary -The primitive accepts generic structured input and emits generic structured output. It must not require a Taslos Tasks database, worker, scheduler, package, runtime path, or private service. +The reference primitive accepts generic test-run bytes and emits generic +structured evidence. It must not require a Taslos Tasks database, worker, +scheduler, package, runtime path, private service, model, provider, or network. +Decision Evidence Protocol and Agent Trajectory Profiler compatibility is by +documented fields, not package imports. + +## Input contract + +An input uses protocol +`opsle.context-firewall.test-run-input/v1` and contains: + +- `operation_id`: caller-supplied stable operation identity; +- `source.id` and optional `source.run_id`; +- `source.raw_evidence_ref`: caller-owned address for raw bytes; +- `process.exit_code`: integer 0 through 255; +- optional supplied `process.duration_ms` and `process.interrupted`; +- one or two uniquely named `stdout`/`stderr` streams encoded as UTF-8 or base64. + +Invalid envelope shape is rejected. Missing semantic provenance is represented +in a packet and triggers escalation when the envelope can still be parsed. + +## Output contract + +The canonical JSON packet contains: + +- `status`: `passed`, `failed`, or `indeterminate`; +- `disposition`: `SUFFICIENT` or `NEEDS_RAW_EVIDENCE`; +- stable `reason_codes` for every insufficiency; +- process status, aggregate counts, failure regions, fatal/timeout/warning data, + and unclassified evidence; +- source, reducer, policy, configuration, retained/suppressed, raw-evidence, and + measurement receipts; +- SHA-256 source-byte, configuration, and semantic-payload identities. + +Object keys are serialized in lexical order with one final newline. Arrays retain +source order. No time, latency, random ID, filesystem state, locale, or ambient +environment value enters canonical output. -## Inputs +## Evidence taxonomy -- `protocol_version`: explicit version. -- `operation_id`: stable idempotency identity. -- `policy_revision`: the exact policy/configuration revision. -- `subject`: vendor-neutral request data required by this concept. -- `evidence`: observable artifacts with provenance. +Source lines have exactly one class: -## Outputs +- `successful_test`, `skipped_test`; +- `failed_test`, `failure_message`, `assertion`, `stack_trace`, + `failure_detail`; +- `fatal_error`, `timeout`, `abnormal_warning`; +- `aggregate_source`, `duration_source`, `structure`, `informational`, `blank`; +- `unclassified` or `unclassified_binary`. -- `status`: accepted, rejected, deferred, or indeterminate. -- `reason`: durable structured reason. -- `changed_entities`: bounded identities, never ambient state. -- `evidence`: provenance-linked receipts. -- `uncertainty`: explicit unknowns. +Derived evidence adds run verdict, process status, aggregate counts, stream +provenance, and reason codes. Matching is structural and strict; keywords inside +otherwise valid test names or explicit notes have no special meaning. ## Invariants -- Information required for safe or correct reasoning is never intentionally suppressed. -- Every reduction is deterministic for a fixed input and policy revision. -- Provenance and raw-output escalation remain available. -- Payload ceilings fail according to an explicit policy. +1. Fixed source bytes, invocation semantics, reducer version, policy revision, + and options produce byte-identical output. +2. Every recognized failed-test region is retained in full in a sufficient + packet. +3. Each source event is counted as retained or suppressed, never neither or + both. +4. Suppression from model context is distinct from raw-evidence destruction. + This reducer destroys nothing. +5. Unclassified, malformed, contradictory, truncated, or under-provenanced + evidence cannot yield a `SUFFICIENT` disposition. +6. A payload ceiling never silently turns omitted critical evidence into a + sufficient packet. +7. `measurements.reduced_bytes` equals the canonical serialized packet length. +8. The input hash binds stream names, lengths, order, and exact bytes. + +## Payload policy + +The deterministic priority order is verdict/process, fatal/timeout, failed test +identity and complete region, aggregates, abnormal warnings, unclassified data, +then repetitive success/structure. + +The default packet suppresses repetitive success and structural lines. If the +packet exceeds `maxOutputBytes`: + +1. warning and unclassified text becomes hash/location references and the packet + requires raw evidence; +2. if complete critical regions still cannot fit, a compact escalation packet + retains failed identities but declares critical evidence omitted; +3. if the compact packet cannot fit, no valid stdout packet is emitted and a + typed `PAYLOAD_CEILING_TOO_SMALL` control error is returned. + +## Raw-evidence escalation + +`NEEDS_RAW_EVIDENCE` applies for unclassified/malformed evidence, contradictions, +interruption, missing process or source identity, missing raw reference, +unexplained nonzero exit, or payload-driven omission. The receipt records whether +a caller reference was supplied. It does not claim that the external artifact +exists, is immutable, or was verified. + +## Idempotency and crash consistency + +The reducer is pure apart from reading CLI input and writing one result. It has no +store and makes no ambient mutation. Repeating an invocation is the idempotency +mechanism. The caller owns durable raw evidence and atomic publication if those +properties are required. -## Failure behavior +## Conformance requirements -Missing required authority or evidence fails closed. Unsupported optional data remains explicit and does not silently widen behavior. Implementations must document idempotency, crash consistency, and raw-evidence escalation. +An implementation conforms only when automated tests cover deterministic replay, +large success reduction, complete single/multi-failure retention, aggregate +correctness, stdout/stderr provenance, exact measurements, ambiguous/malformed +input, payload boundaries, typed escalation, and absence of generated semantic +time/randomness. ## Versioning -Breaking semantic changes require a new protocol version. New optional fields require evidence that they affect a real decision. +Breaking input, output, classification, priority, hash-framing, or escalation +semantics require a new protocol or policy version. Optional fields require +evidence of decision value before admission. diff --git a/bin/context-firewall.js b/bin/context-firewall.js new file mode 100755 index 0000000..d3fef8f --- /dev/null +++ b/bin/context-firewall.js @@ -0,0 +1,68 @@ +#!/usr/bin/env node +import { readFile } from 'node:fs/promises'; +import process from 'node:process'; +import { conformanceReport } from '../fixtures/corpus.js'; +import { + InputError, + PayloadCeilingError, + canonicalJson, + reduceTestRun, + serializePacket, +} from '../src/reducer.js'; + +function usage() { + return 'usage: context-firewall reduce [--input PATH|-] [--max-bytes N]\n context-firewall conformance\n'; +} + +function parseReduceArgs(args) { + let input = '-'; + let maxOutputBytes = null; + for (let index = 0; index < args.length; index += 1) { + if (args[index] === '--input' && args[index + 1]) input = args[++index]; + else if (args[index] === '--max-bytes' && args[index + 1]) { + maxOutputBytes = Number(args[++index]); + } else throw new InputError(`unknown or incomplete argument: ${args[index]}`); + } + return { input, maxOutputBytes }; +} + +async function readInput(path) { + if (path !== '-') return readFile(path); + const chunks = []; + for await (const chunk of process.stdin) chunks.push(chunk); + return Buffer.concat(chunks); +} + +async function main() { + const [command, ...args] = process.argv.slice(2); + if (command === 'conformance') { + if (args.length > 0) throw new InputError('conformance accepts no arguments'); + const report = conformanceReport(); + process.stdout.write(`${canonicalJson(report)}\n`); + process.exitCode = report.conformance === 'PASS' ? 0 : 1; + return; + } + if (command !== 'reduce') { + process.stderr.write(usage()); + process.exitCode = 2; + return; + } + const options = parseReduceArgs(args); + const bytes = await readInput(options.input); + let input; + try { + input = JSON.parse(bytes.toString('utf8')); + } catch { + throw new InputError('input must be valid JSON'); + } + const packet = reduceTestRun(input, { maxOutputBytes: options.maxOutputBytes }); + process.stdout.write(serializePacket(packet)); +} + +main().catch((error) => { + const body = error instanceof PayloadCeilingError + ? { code: error.code, ...error.details } + : { code: error.code ?? 'INTERNAL_ERROR', message: error.message }; + process.stderr.write(`${canonicalJson(body)}\n`); + process.exitCode = error instanceof InputError || error instanceof PayloadCeilingError ? 2 : 1; +}); diff --git a/examples/test-run-input.json b/examples/test-run-input.json new file mode 100644 index 0000000..56e7585 --- /dev/null +++ b/examples/test-run-input.json @@ -0,0 +1,21 @@ +{ + "protocol_version": "opsle.context-firewall.test-run-input/v1", + "operation_id": "op-public-example", + "source": { + "id": "synthetic-suite", + "run_id": "run-public-example", + "raw_evidence_ref": "artifact://public/run.tap" + }, + "process": { + "exit_code": 1, + "duration_ms": 12.5, + "interrupted": false + }, + "streams": [ + { + "name": "stdout", + "encoding": "utf8", + "data": "TAP version 13\nnot ok 1 - adds\n message: expected 2\n expected: 2\n actual: 3\n1..1\n# tests 1\n# pass 0\n# fail 1\n# skipped 0\n# duration_ms 12.5\n" + } + ] +} diff --git a/fixtures/README.md b/fixtures/README.md new file mode 100644 index 0000000..4a4a972 --- /dev/null +++ b/fixtures/README.md @@ -0,0 +1,14 @@ +# Synthetic conformance corpus + +`corpus.js` defines 30 deterministic public-safe fixtures. Large and repetitive +transcripts are generated from fixed recipes so the repository does not contain +unnecessary generated logs. Every fixture has an expected verdict, retained +evidence, escalation state, or ceiling failure. + +The corpus covers normal success and skipping; single, mixed, and multiple +failures; assertion and stack details; runner crash, unexplained exit, timeout, +and interruption; stdout/stderr, ANSI, Unicode, long lines, missing newline, +repetition, malformed output, misleading keywords; ceiling tiers; and supplied +or missing provenance. + +This is conformance evidence only, not an experimental dataset or benchmark. diff --git a/fixtures/corpus.js b/fixtures/corpus.js new file mode 100644 index 0000000..cb7c17e --- /dev/null +++ b/fixtures/corpus.js @@ -0,0 +1,179 @@ +import { INPUT_PROTOCOL, PayloadCeilingError, reduceTestRun, serializePacket } from '../src/reducer.js'; + +function tap({ passed = [], failed = [], skipped = [], extras = [], finalNewline = true }) { + const lines = ['TAP version 13']; + let index = 1; + for (const name of passed) lines.push(`ok ${index++} - ${name}`); + for (const name of skipped) lines.push(`ok ${index++} - ${name} # SKIP synthetic`); + for (const failure of failed) { + lines.push(`not ok ${index++} - ${failure.name}`); + lines.push(...(failure.details ?? []).map((detail) => ` ${detail}`)); + } + lines.push(...extras); + lines.push(`1..${passed.length + failed.length + skipped.length}`); + lines.push(`# tests ${passed.length + failed.length + skipped.length}`); + lines.push(`# pass ${passed.length}`); + lines.push(`# fail ${failed.length}`); + lines.push(`# skipped ${skipped.length}`); + lines.push('# duration_ms 12.5'); + return `${lines.join('\n')}${finalNewline ? '\n' : ''}`; +} + +function invocation({ + stdout = '', stderr = '', exitCode = 0, interrupted = false, + source = true, rawRef = true, operationId = true, durationMs = 12.5, +} = {}) { + const streams = [{ name: 'stdout', encoding: 'utf8', data: stdout }]; + if (stderr !== '') streams.push({ name: 'stderr', encoding: 'utf8', data: stderr }); + return { + protocol_version: INPUT_PROTOCOL, + ...(operationId ? { operation_id: 'op-synthetic-001' } : {}), + source: { + ...(source ? { id: 'synthetic-suite', run_id: 'run-001' } : {}), + ...(rawRef ? { raw_evidence_ref: 'artifact://synthetic/run-001.tap' } : {}), + }, + process: { exit_code: exitCode, duration_ms: durationMs, interrupted }, + streams, + }; +} + +const manyPasses = Array.from({ length: 1500 }, (_, index) => `case-${String(index + 1).padStart(4, '0')}`); +const repetitivePasses = Array.from({ length: 400 }, () => 'repetitive-success'); +const longName = `very-long-${'x'.repeat(40_000)}`; +const longFailure = `message: ${'critical'.repeat(4_000)}`; + +export const corpus = Object.freeze([ + { name: 'normal/small-all-pass', input: invocation({ stdout: tap({ passed: ['alpha', 'beta'] }) }), expected: { status: 'passed', disposition: 'SUFFICIENT', counts: [2, 0, 0] } }, + { name: 'normal/large-all-pass', input: invocation({ stdout: tap({ passed: manyPasses }) }), expected: { status: 'passed', disposition: 'SUFFICIENT', counts: [1500, 0, 0], substantialReduction: true } }, + { name: 'normal/repetitive-success', input: invocation({ stdout: tap({ passed: repetitivePasses }) }), expected: { status: 'passed', counts: [400, 0, 0], substantialReduction: true } }, + { name: 'normal/skipped-tests', input: invocation({ stdout: tap({ passed: ['runs'], skipped: ['not-applicable', 'platform-only'] }) }), expected: { status: 'passed', counts: [1, 0, 2] } }, + { name: 'failure/one', input: invocation({ stdout: tap({ failed: [{ name: 'adds values', details: ['message: expected two', 'expected: 2', 'actual: 3'] }] }), exitCode: 1 }), expected: { status: 'failed', failures: 1, contains: ['adds values', 'expected two'] } }, + { name: 'failure/several', input: invocation({ stdout: tap({ failed: [{ name: 'first', details: ['message: first broke'] }, { name: 'second', details: ['message: second broke'] }, { name: 'third', details: ['message: third broke'] }] }), exitCode: 1 }), expected: { status: 'failed', failures: 3, contains: ['first broke', 'second broke', 'third broke'] } }, + { name: 'failure/mixed-pass-fail', input: invocation({ stdout: tap({ passed: ['green-a', 'green-b'], failed: [{ name: 'red', details: ['message: broken'] }] }), exitCode: 1 }), expected: { status: 'failed', counts: [2, 1, 0], failures: 1 } }, + { name: 'failure/assertion-difference', input: invocation({ stdout: tap({ failed: [{ name: 'diffs objects', details: ['operator: deepStrictEqual', 'expected: {a: 1}', 'actual: {a: 2}', 'diff: -1 +2'] }] }), exitCode: 1 }), expected: { status: 'failed', contains: ['deepStrictEqual', 'diff: -1 +2'] } }, + { name: 'failure/stack-trace', input: invocation({ stdout: tap({ failed: [{ name: 'throws', details: ['---', 'error: synthetic failure', 'code: ERR_SYNTHETIC', 'stack: |', ' TypeError: synthetic failure', ' at subject (file:///public/example.js:2:3)', ' at test (file:///public/test.js:4:5)', '...'] }] }), exitCode: 1 }), expected: { status: 'failed', contains: ['TypeError: synthetic failure', 'file:///public/test.js:4:5'] } }, + { name: 'failure/independent-regions', input: invocation({ stdout: tap({ passed: ['between'], failed: [{ name: 'region-a', details: ['message: A'] }, { name: 'region-b', details: ['message: B'] }] }), exitCode: 1 }), expected: { status: 'failed', failures: 2, contains: ['region-a', 'region-b'] } }, + { name: 'process/runner-crash', input: invocation({ stdout: tap({ passed: ['before-crash'] }), stderr: 'FATAL: runner crashed synthetically\n', exitCode: 2 }), expected: { status: 'failed', fatal: 1, contains: ['runner crashed synthetically'] } }, + { name: 'process/nonzero-unrecognized', input: invocation({ stdout: '# note: process stopped without a test failure\n', exitCode: 2 }), expected: { status: 'failed', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'NONZERO_EXIT_WITHOUT_RECOGNIZED_FAILURE' } }, + { name: 'process/timeout', input: invocation({ stdout: 'TIMEOUT: suite exceeded 5000ms\n', exitCode: 124 }), expected: { status: 'failed', timeout: 1, contains: ['suite exceeded 5000ms'] } }, + { name: 'process/interrupted-truncated', input: invocation({ stdout: 'TAP version 13\nnot ok 1 - incomplete\n message: cut off', exitCode: 1, interrupted: true }), expected: { status: 'failed', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'INTERRUPTED_OUTPUT', contains: ['cut off'] } }, + { name: 'edge/mixed-stdout-stderr', input: invocation({ stdout: tap({ passed: ['stdout-case'] }), stderr: 'WARNING: synthetic stderr warning\n' }), expected: { status: 'passed', warning: 1, streams: 2 } }, + { name: 'edge/ansi', input: invocation({ stdout: tap({ failed: [{ name: '\u001b[31mcolored failure\u001b[0m', details: ['message: \u001b[31mred\u001b[0m'] }] }), exitCode: 1 }), expected: { status: 'failed', failures: 1, contains: ['colored failure', 'red'] } }, + { name: 'edge/unicode', input: invocation({ stdout: tap({ failed: [{ name: 'unicode-λ-🧪', details: ['message: café ≠ чай'] }] }), exitCode: 1 }), expected: { status: 'failed', contains: ['unicode-λ-🧪', 'café ≠ чай'] } }, + { name: 'edge/very-long-line', input: invocation({ stdout: tap({ passed: [longName] }) }), expected: { status: 'passed', substantialReduction: true } }, + { name: 'edge/missing-final-newline', input: invocation({ stdout: tap({ passed: ['no-newline'] , finalNewline: false }) }), expected: { status: 'passed', counts: [1, 0, 0] } }, + { name: 'edge/repeated-identical-messages', input: invocation({ stdout: tap({ passed: ['still-green'], extras: Array.from({ length: 300 }, () => '# note: identical heartbeat') }) }), expected: { status: 'passed', substantialReduction: true } }, + { name: 'edge/malformed-output', input: invocation({ stdout: 'this is not TAP\nnot okay maybe FAIL\n' }), expected: { status: 'indeterminate', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'UNCLASSIFIED_EVIDENCE' } }, + { name: 'edge/superficial-pass-fail', input: invocation({ stdout: tap({ passed: ['literal PASS FAIL error warning text'], extras: ['# note: PASS FAIL error warning are prose'] }) }), expected: { status: 'passed', failures: 0, warning: 0 } }, + { name: 'payload/comfortably-above', input: invocation({ stdout: tap({ failed: [{ name: 'bounded', details: ['message: retained'] }] }), exitCode: 1 }), options: { maxOutputBytes: 20_000 }, expected: { status: 'failed', failures: 1, payloadAffected: false } }, + { name: 'payload/exact-boundary', input: invocation({ stdout: tap({ passed: manyPasses }) }), ceilingMode: 'exact', expected: { status: 'passed', payloadAffected: false, substantialReduction: true } }, + { name: 'payload/critical-too-large', input: invocation({ stdout: tap({ failed: [{ name: 'oversized-critical', details: [longFailure] }] }), exitCode: 1 }), options: { maxOutputBytes: 3_500 }, expected: { status: 'failed', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'PAYLOAD_LIMIT_CRITICAL_EVIDENCE_EXCEEDED', payloadAffected: true } }, + { name: 'payload/safe-packet-impossible', input: invocation({ stdout: tap({ failed: [{ name: 'cannot-fit', details: [longFailure] }] }), exitCode: 1 }), options: { maxOutputBytes: 64 }, expected: { ceilingError: true } }, + { name: 'provenance/raw-reference-present', input: invocation({ stdout: tap({ passed: ['addressable'] }) }), expected: { status: 'passed', disposition: 'SUFFICIENT', rawRef: true } }, + { name: 'provenance/raw-reference-absent', input: invocation({ stdout: tap({ passed: ['not-addressable'] }), rawRef: false }), expected: { status: 'passed', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'RAW_EVIDENCE_REFERENCE_MISSING', rawRef: false } }, + { name: 'provenance/source-present', input: invocation({ stdout: tap({ passed: ['identified'] }) }), expected: { status: 'passed', source: true } }, + { name: 'provenance/source-missing', input: invocation({ stdout: tap({ passed: ['anonymous'] }), source: false }), expected: { status: 'indeterminate', disposition: 'NEEDS_RAW_EVIDENCE', reason: 'SOURCE_IDENTITY_MISSING', source: false } }, +]); + +function exactCeiling(input) { + let ceiling = 100_000; + for (let iteration = 0; iteration < 20; iteration += 1) { + const packet = reduceTestRun(input, { maxOutputBytes: ceiling }); + const size = serializePacket(packet).length; + if (size === ceiling) return ceiling; + ceiling = size; + } + throw new Error('exact ceiling did not converge'); +} + +export function executeFixture(fixture) { + const options = fixture.ceilingMode === 'exact' + ? { maxOutputBytes: exactCeiling(fixture.input) } + : (fixture.options ?? {}); + try { + const packet = reduceTestRun(fixture.input, options); + const bytes = serializePacket(packet); + const evidence = packet.decision_evidence; + const expected = fixture.expected; + const text = bytes.toString('utf8'); + const checks = []; + if (expected.ceilingError) checks.push(false); + if (expected.status) checks.push(evidence.status === expected.status); + if (expected.disposition) checks.push(evidence.disposition === expected.disposition); + if (expected.counts) checks.push( + evidence.counts.passed === expected.counts[0] + && evidence.counts.failed === expected.counts[1] + && evidence.counts.skipped === expected.counts[2], + ); + if (expected.failures != null) checks.push(evidence.failures.length === expected.failures); + if (expected.fatal != null) checks.push(evidence.fatal_errors.length === expected.fatal); + if (expected.timeout != null) checks.push(evidence.timeouts.length === expected.timeout); + if (expected.warning != null) checks.push(evidence.warnings.length === expected.warning); + if (expected.reason) checks.push(evidence.reason_codes.includes(expected.reason)); + if (expected.contains) checks.push(expected.contains.every((needle) => text.includes(needle))); + if (expected.payloadAffected != null) checks.push(packet.receipt.payload_limit.affected === expected.payloadAffected); + if (expected.streams != null) checks.push(packet.receipt.source.streams.length === expected.streams); + if (expected.rawRef != null) checks.push(packet.receipt.raw_evidence.escalation_available === expected.rawRef); + if (expected.source != null) checks.push(Boolean(packet.receipt.source.id) === expected.source); + if (expected.substantialReduction) checks.push( + packet.receipt.measurements.reduced_bytes < packet.receipt.measurements.original_bytes * 0.5, + ); + if (options.maxOutputBytes != null) checks.push(bytes.length <= options.maxOutputBytes); + checks.push(packet.receipt.measurements.reduced_bytes === bytes.length); + checks.push(packet.receipt.measurements.original_bytes === fixture.input.streams.reduce( + (sum, stream) => sum + Buffer.byteLength(stream.data, stream.encoding === 'base64' ? 'base64' : 'utf8'), 0, + )); + return { packet, options, pass: checks.every(Boolean) }; + } catch (error) { + if (error instanceof PayloadCeilingError && fixture.expected.ceilingError) { + return { error, options, pass: true }; + } + return { error, options, pass: false }; + } +} + +export function conformanceReport() { + const fixtures = corpus.map((fixture) => { + const result = executeFixture(fixture); + if (result.error) { + return { + conformance: result.pass ? 'PASS' : 'FAIL', + escalation_required: true, + fixture: fixture.name, + raw_bytes: fixture.input.streams.reduce( + (sum, stream) => sum + Buffer.byteLength(stream.data, stream.encoding === 'base64' ? 'base64' : 'utf8'), + 0, + ), + reduced_bytes: null, + reduction_percentage: null, + verdict_preserved: null, + }; + } + const packet = result.packet; + return { + conformance: result.pass ? 'PASS' : 'FAIL', + critical_evidence_preserved: packet.decision_evidence.disposition === 'SUFFICIENT' + || packet.receipt.raw_evidence.escalation_required, + escalation_required: packet.receipt.raw_evidence.escalation_required, + fixture: fixture.name, + raw_bytes: packet.receipt.measurements.original_bytes, + reduced_bytes: packet.receipt.measurements.reduced_bytes, + reduction_percentage: packet.receipt.measurements.original_bytes === 0 + ? null + : Number(( + (packet.receipt.measurements.original_bytes - packet.receipt.measurements.reduced_bytes) + / packet.receipt.measurements.original_bytes + * 100 + ).toFixed(2)), + verdict_preserved: fixture.expected.status == null + ? null + : packet.decision_evidence.status === fixture.expected.status, + }; + }); + return { + conformance: fixtures.every((fixture) => fixture.conformance === 'PASS') ? 'PASS' : 'FAIL', + fixture_count: fixtures.length, + fixtures, + protocol_version: 'opsle.context-firewall.conformance/v1', + }; +} diff --git a/package.json b/package.json new file mode 100644 index 0000000..66c7f12 --- /dev/null +++ b/package.json @@ -0,0 +1,16 @@ +{ + "name": "@opsle/context-firewall", + "version": "0.2.0", + "private": true, + "type": "module", + "bin": { + "context-firewall": "./bin/context-firewall.js" + }, + "scripts": { + "test": "node --test", + "conformance": "node ./bin/context-firewall.js conformance" + }, + "engines": { + "node": ">=20" + } +} diff --git a/src/README.md b/src/README.md index 40f86b5..d53ea44 100644 --- a/src/README.md +++ b/src/README.md @@ -1,3 +1,11 @@ # Reference surface -Implementation is deliberately deferred until the theory and interface justify code. Begin with THEORY.md, SPEC.md, and BENCHMARK.md. +`reducer.js` exports the dependency-free core: + +- `normalizeInvocation(value)` validates and decodes the public input envelope; +- `reduceTestRun(value, options)` returns a deterministic evidence packet; +- `serializePacket(packet)` returns canonical UTF-8 JSON with a final newline; +- `PayloadCeilingError` represents a ceiling too small for any safe packet. + +The core performs no I/O, model calls, network calls, persistence, or host +integration. diff --git a/src/reducer.js b/src/reducer.js new file mode 100644 index 0000000..2be1c82 --- /dev/null +++ b/src/reducer.js @@ -0,0 +1,518 @@ +import { createHash } from 'node:crypto'; +import { TextDecoder } from 'node:util'; + +export const INPUT_PROTOCOL = 'opsle.context-firewall.test-run-input/v1'; +export const PACKET_PROTOCOL = 'opsle.context-firewall.evidence-packet/v1'; +export const REDUCER_NAME = '@opsle/context-firewall/test-output'; +export const REDUCER_VERSION = '0.2.0'; +export const POLICY_REVISION = 'tap-subset-policy/v1'; + +const ANSI_PATTERN = /[\u001b\u009b][[\]()#;?]*(?:(?:(?:[a-zA-Z\d]*(?:;[-a-zA-Z\d/#&.:=?%@~_]+)*)?\u0007)|(?:(?:\d{1,4}(?:[;:]\d{0,4})*)?[\dA-PR-TZcf-nq-uy=><~]))/g; +const utf8Decoder = new TextDecoder('utf-8', { fatal: true }); +const streamOrder = new Map([['stdout', 0], ['stderr', 1]]); + +export class InputError extends Error { + constructor(message) { + super(message); + this.name = 'InputError'; + this.code = 'INVALID_INPUT'; + } +} + +export class PayloadCeilingError extends Error { + constructor(details) { + super('payload ceiling is too small for a safe evidence packet'); + this.name = 'PayloadCeilingError'; + this.code = 'PAYLOAD_CEILING_TOO_SMALL'; + this.details = details; + } +} + +function isPlainObject(value) { + return value !== null && typeof value === 'object' && !Array.isArray(value); +} + +function canonicalize(value) { + if (Array.isArray(value)) return value.map(canonicalize); + if (!isPlainObject(value)) return value; + return Object.fromEntries( + Object.keys(value).sort().map((key) => [key, canonicalize(value[key])]), + ); +} + +export function canonicalJson(value) { + return JSON.stringify(canonicalize(value)); +} + +export function sha256(value) { + const digest = createHash('sha256'); + digest.update(value); + return `sha256:${digest.digest('hex')}`; +} + +function decodeBase64(value) { + if (value.length % 4 !== 0 || !/^(?:[A-Za-z0-9+/]{4})*(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?$/.test(value)) { + throw new InputError('stream base64 data is invalid'); + } + return Buffer.from(value, 'base64'); +} + +export function normalizeInvocation(value) { + if (!isPlainObject(value)) throw new InputError('input object required'); + if (value.protocol_version !== INPUT_PROTOCOL) { + throw new InputError(`protocol_version must be ${INPUT_PROTOCOL}`); + } + if (!Array.isArray(value.streams) || value.streams.length === 0) { + throw new InputError('at least one stream is required'); + } + + const seen = new Set(); + const streams = value.streams.map((stream) => { + if (!isPlainObject(stream) || !streamOrder.has(stream.name)) { + throw new InputError('stream name must be stdout or stderr'); + } + if (seen.has(stream.name)) throw new InputError(`duplicate ${stream.name} stream`); + seen.add(stream.name); + if (typeof stream.data !== 'string') throw new InputError('stream data must be a string'); + const encoding = stream.encoding ?? 'utf8'; + if (!['utf8', 'base64'].includes(encoding)) { + throw new InputError('stream encoding must be utf8 or base64'); + } + return { + name: stream.name, + bytes: encoding === 'base64' ? decodeBase64(stream.data) : Buffer.from(stream.data, 'utf8'), + }; + }).sort((left, right) => streamOrder.get(left.name) - streamOrder.get(right.name)); + + const source = isPlainObject(value.source) ? value.source : {}; + const process = isPlainObject(value.process) ? value.process : {}; + if (process.exit_code != null && (!Number.isInteger(process.exit_code) || process.exit_code < 0 || process.exit_code > 255)) { + throw new InputError('process.exit_code must be an integer from 0 through 255 or null'); + } + if (process.duration_ms != null && (typeof process.duration_ms !== 'number' || !Number.isFinite(process.duration_ms) || process.duration_ms < 0)) { + throw new InputError('process.duration_ms must be a nonnegative finite number or null'); + } + if (process.interrupted != null && typeof process.interrupted !== 'boolean') { + throw new InputError('process.interrupted must be boolean when supplied'); + } + + return { + protocolVersion: value.protocol_version, + operationId: typeof value.operation_id === 'string' && value.operation_id.length > 0 + ? value.operation_id + : null, + source: { + id: typeof source.id === 'string' && source.id.length > 0 ? source.id : null, + runId: typeof source.run_id === 'string' && source.run_id.length > 0 ? source.run_id : null, + rawEvidenceRef: typeof source.raw_evidence_ref === 'string' && source.raw_evidence_ref.length > 0 + ? source.raw_evidence_ref + : null, + }, + process: { + exitCode: process.exit_code ?? null, + durationMs: process.duration_ms ?? null, + interrupted: process.interrupted ?? false, + }, + streams, + }; +} + +function streamDigest(streams) { + const digest = createHash('sha256'); + for (const stream of streams) { + digest.update(Buffer.from(`${stream.name}\0${stream.bytes.length}\0`, 'utf8')); + digest.update(stream.bytes); + digest.update(Buffer.from('\0', 'utf8')); + } + return `sha256:${digest.digest('hex')}`; +} + +function splitLines(text) { + if (text.length === 0) return []; + const lines = text.split('\n'); + if (text.endsWith('\n')) lines.pop(); + return lines; +} + +function lineEvidence(line, category, includeText = true) { + const value = { + category, + line: line.line, + sha256: sha256(Buffer.from(line.text, 'utf8')), + stream: line.stream, + }; + if (includeText) value.text = line.text; + else value.byte_count = Buffer.byteLength(line.text, 'utf8'); + return value; +} + +function classifyFailureDetail(clean) { + if (/^\s*(?:operator|expected|actual|diff|code):/i.test(clean)) return 'assertion'; + if (/^\s*(?:[A-Za-z]*Error:|at\s+|.*\([^()]+:\d+:\d+\)\s*$)/.test(clean)) return 'stack_trace'; + if (/^\s*(?:error|message|cause):/i.test(clean)) return 'failure_message'; + return 'failure_detail'; +} + +function parseTranscript(streams) { + const lines = []; + const failures = []; + const fatalErrors = []; + const timeouts = []; + const warnings = []; + const unclassified = []; + const summaryValues = new Map(); + const summaryConflicts = []; + const observed = { passed: 0, failed: 0, skipped: 0 }; + let malformedUtf8 = false; + + for (const stream of streams) { + let text; + try { + text = utf8Decoder.decode(stream.bytes); + } catch { + malformedUtf8 = true; + if (stream.bytes.length > 0) { + const line = { + stream: stream.name, + line: 1, + text: `[non-UTF-8 ${stream.bytes.length} bytes; ${sha256(stream.bytes)}]`, + kind: 'unclassified_binary', + }; + lines.push(line); + unclassified.push(line); + } + continue; + } + + let currentFailure = null; + for (const [index, rawLine] of splitLines(text).entries()) { + const line = { stream: stream.name, line: index + 1, text: rawLine }; + const clean = rawLine.replace(/\r$/, '').replace(ANSI_PATTERN, ''); + const marker = clean.match(/^\s*(not ok|ok)\b(?:\s+\d+)?(?:\s*-\s*)?(.*)$/); + const summary = clean.match(/^\s*#\s*(tests|pass|fail|skipped)\s+(\d+)\s*$/); + + if (marker) { + currentFailure = null; + const directive = marker[2].match(/\s+#\s*(SKIP|TODO)\b.*$/i); + const identity = marker[2].replace(/\s+#\s*(?:SKIP|TODO)\b.*$/i, '').trim() || '(unnamed test)'; + if (marker[1] === 'ok' && directive?.[1].toUpperCase() === 'SKIP') { + line.kind = 'skipped_test'; + observed.skipped += 1; + } else if (marker[1] === 'ok') { + line.kind = 'successful_test'; + observed.passed += 1; + } else if (directive?.[1].toUpperCase() === 'TODO') { + line.kind = 'unclassified'; + unclassified.push(line); + } else { + line.kind = 'failed_test'; + observed.failed += 1; + currentFailure = { identity, header: line, details: [] }; + failures.push(currentFailure); + } + } else if (summary) { + currentFailure = null; + line.kind = 'aggregate_source'; + const key = summary[1] === 'pass' ? 'passed' + : summary[1] === 'fail' ? 'failed' + : summary[1]; + const value = Number(summary[2]); + if (summaryValues.has(key) && summaryValues.get(key) !== value) { + summaryConflicts.push(key); + } + summaryValues.set(key, value); + } else if (currentFailure && !/^\s*(?:TAP version \d+|1\.\.\d+|#\s*Subtest:.*)\s*$/.test(clean)) { + line.kind = classifyFailureDetail(clean); + currentFailure.details.push(line); + } else if (/^\s*(?:TAP version \d+|1\.\.\d+|#\s*Subtest:.*|---|\.\.\.)\s*$/.test(clean)) { + currentFailure = null; + line.kind = 'structure'; + } else { + const fatal = clean.match(/^\s*(?:#\s*)?(?:FATAL|RUNNER CRASH|UNCAUGHT):\s*(.+)$/); + const timeout = clean.match(/^\s*(?:#\s*)?TIMEOUT:\s*(.+)$/); + const warning = clean.match(/^\s*(?:#\s*)?(?:WARNING|WARN):\s*(.+)$/); + if (fatal) { + line.kind = 'fatal_error'; + fatalErrors.push(line); + } else if (timeout) { + line.kind = 'timeout'; + timeouts.push(line); + } else if (warning) { + line.kind = 'abnormal_warning'; + warnings.push(line); + } else if (/^\s*#\s*(?:duration_ms\s+\d+(?:\.\d+)?|note:\s*.*)\s*$/.test(clean)) { + line.kind = clean.includes('duration_ms') ? 'duration_source' : 'informational'; + } else if (/^\s*$/.test(clean)) { + line.kind = 'blank'; + } else { + line.kind = 'unclassified'; + unclassified.push(line); + } + } + lines.push(line); + } + } + + const counts = { + passed: summaryValues.get('passed') ?? observed.passed, + failed: summaryValues.get('failed') ?? observed.failed, + skipped: summaryValues.get('skipped') ?? observed.skipped, + total: summaryValues.get('tests') + ?? ((summaryValues.has('passed') || summaryValues.has('failed') || summaryValues.has('skipped')) + ? (summaryValues.get('passed') ?? 0) + (summaryValues.get('failed') ?? 0) + (summaryValues.get('skipped') ?? 0) + : observed.passed + observed.failed + observed.skipped), + }; + const contradictions = [...summaryConflicts]; + for (const key of ['passed', 'failed', 'skipped']) { + if (summaryValues.has(key) && summaryValues.get(key) !== observed[key]) contradictions.push(key); + } + if (summaryValues.has('tests') && summaryValues.get('tests') !== observed.passed + observed.failed + observed.skipped) { + contradictions.push('total'); + } + + return { + lines, + failures, + fatalErrors, + timeouts, + warnings, + unclassified, + counts, + contradictions: [...new Set(contradictions)].sort(), + malformedUtf8, + }; +} + +function categoryCounts(lines) { + const counts = {}; + for (const line of lines) counts[line.kind] = (counts[line.kind] ?? 0) + 1; + return Object.fromEntries(Object.entries(counts).sort(([left], [right]) => left.localeCompare(right))); +} + +function addReason(reasons, reason) { + if (!reasons.includes(reason)) reasons.push(reason); +} + +function evidenceStatus(invocation, parsed, reasons) { + const recognizedFailure = parsed.failures.length > 0 + || parsed.fatalErrors.length > 0 + || parsed.timeouts.length > 0 + || parsed.counts.failed > 0; + const nonzero = invocation.process.exitCode != null && invocation.process.exitCode !== 0; + + if (parsed.unclassified.length > 0) addReason(reasons, 'UNCLASSIFIED_EVIDENCE'); + if (parsed.malformedUtf8) addReason(reasons, 'MALFORMED_UTF8'); + if (parsed.contradictions.length > 0) addReason(reasons, 'AGGREGATE_CONTRADICTION'); + if (invocation.process.interrupted) addReason(reasons, 'INTERRUPTED_OUTPUT'); + if (invocation.process.exitCode == null) addReason(reasons, 'EXIT_STATUS_MISSING'); + if (nonzero && !recognizedFailure) addReason(reasons, 'NONZERO_EXIT_WITHOUT_RECOGNIZED_FAILURE'); + if (!invocation.operationId) addReason(reasons, 'OPERATION_ID_MISSING'); + if (!invocation.source.id) addReason(reasons, 'SOURCE_IDENTITY_MISSING'); + if (!invocation.source.rawEvidenceRef) addReason(reasons, 'RAW_EVIDENCE_REFERENCE_MISSING'); + + if (recognizedFailure || nonzero) return 'failed'; + if (reasons.some((reason) => !['RAW_EVIDENCE_REFERENCE_MISSING'].includes(reason))) return 'indeterminate'; + return 'passed'; +} + +function measurements(originalBytes, originalEvents, retainedLines, packetBytes) { + return { + original_bytes: originalBytes, + original_event_count: originalEvents, + reduced_bytes: packetBytes, + retained_evidence_count: retainedLines, + suppressed_evidence_count: originalEvents - retainedLines, + }; +} + +function finalizePacket(packet, originalBytes, originalEvents, retainedLines) { + let reducedBytes = 0; + for (let iteration = 0; iteration < 12; iteration += 1) { + packet.receipt.measurements = measurements( + originalBytes, + originalEvents, + retainedLines, + reducedBytes, + ); + const next = Buffer.byteLength(`${canonicalJson(packet)}\n`, 'utf8'); + if (next === reducedBytes) return packet; + reducedBytes = next; + } + throw new Error('reduced byte measurement did not converge'); +} + +function buildPacket({ invocation, parsed, configuration, baseReasons, mode }) { + const includeWarnings = mode === 'full'; + const includeUnclassified = mode === 'full'; + const includeCriticalText = mode !== 'compact'; + const reasons = [...baseReasons]; + const payloadAffected = mode !== 'full'; + if (mode === 'references') addReason(reasons, 'PAYLOAD_LIMIT_OMITTED_NONCRITICAL_TEXT'); + if (mode === 'compact') addReason(reasons, 'PAYLOAD_LIMIT_CRITICAL_EVIDENCE_EXCEEDED'); + reasons.sort(); + + const failurePackets = parsed.failures.map((failure) => ({ + identity: failure.identity, + ...(includeCriticalText ? { + header: lineEvidence(failure.header, 'failed_test'), + details: failure.details.map((line) => lineEvidence(line, line.kind)), + } : {}), + })); + const fatalPackets = includeCriticalText + ? parsed.fatalErrors.map((line) => lineEvidence(line, 'fatal_error')) + : []; + const timeoutPackets = includeCriticalText + ? parsed.timeouts.map((line) => lineEvidence(line, 'timeout')) + : []; + const warningPackets = parsed.warnings.map((line) => lineEvidence(line, 'abnormal_warning', includeWarnings)); + const unclassifiedPackets = parsed.unclassified.map((line) => lineEvidence(line, line.kind, includeUnclassified)); + + const retainedKinds = new Set(); + if (includeCriticalText) { + for (const failure of parsed.failures) { + retainedKinds.add(failure.header); + for (const detail of failure.details) retainedKinds.add(detail); + } + for (const line of [...parsed.fatalErrors, ...parsed.timeouts]) retainedKinds.add(line); + } + if (includeWarnings) for (const line of parsed.warnings) retainedKinds.add(line); + if (includeUnclassified) for (const line of parsed.unclassified) retainedKinds.add(line); + const retainedLines = parsed.lines.filter((line) => retainedKinds.has(line)); + const suppressedLines = parsed.lines.filter((line) => !retainedKinds.has(line)); + const retainedCategories = [ + 'aggregate_counts', + 'process_status', + 'run_verdict', + 'stream_provenance', + ...Object.keys(categoryCounts(retainedLines)), + ].sort(); + + const decisionEvidence = { + counts: parsed.counts, + disposition: reasons.length === 0 ? 'SUFFICIENT' : 'NEEDS_RAW_EVIDENCE', + failures: failurePackets, + fatal_errors: fatalPackets, + process: { + duration_ms: invocation.process.durationMs, + exit_code: invocation.process.exitCode, + interrupted: invocation.process.interrupted, + }, + reason_codes: reasons, + source: { + id: invocation.source.id, + run_id: invocation.source.runId, + }, + status: evidenceStatus(invocation, parsed, []), + timeouts: timeoutPackets, + unclassified_evidence: unclassifiedPackets, + warnings: warningPackets, + }; + + const originalBytes = invocation.streams.reduce((sum, stream) => sum + stream.bytes.length, 0); + const inputHash = streamDigest(invocation.streams); + const semanticHash = sha256(Buffer.from(canonicalJson(decisionEvidence), 'utf8')); + const packet = { + decision_evidence: decisionEvidence, + operation_id: invocation.operationId, + protocol_version: PACKET_PROTOCOL, + receipt: { + configuration: { + identity: configuration.identity, + max_output_bytes: configuration.maxOutputBytes, + policy_revision: POLICY_REVISION, + }, + input_hash: inputHash, + measurements: {}, + payload_limit: { + affected: payloadAffected, + honored: true, + requested_bytes: configuration.maxOutputBytes, + }, + raw_evidence: { + destroyed_by_reducer: false, + escalation_available: Boolean(invocation.source.rawEvidenceRef), + escalation_required: reasons.length > 0, + reference: invocation.source.rawEvidenceRef, + source_evidence_disposition: invocation.source.rawEvidenceRef + ? 'CALLER_REFERENCE_SUPPLIED' + : 'PRESERVATION_UNCONFIRMED', + suppressed_from_model_context: suppressedLines.length > 0, + }, + receipt_version: 1, + reducer: { + name: REDUCER_NAME, + version: REDUCER_VERSION, + }, + reduction_complete: reasons.length === 0, + retained: { + categories: retainedCategories, + event_count: retainedLines.length, + }, + semantic_payload_hash: semanticHash, + source: { + id: invocation.source.id, + protocol_version: invocation.protocolVersion, + run_id: invocation.source.runId, + streams: invocation.streams.map((stream) => ({ + byte_count: stream.bytes.length, + name: stream.name, + })), + }, + suppressed: { + categories: categoryCounts(suppressedLines), + event_count: suppressedLines.length, + }, + unclassified_evidence_present: parsed.unclassified.length > 0, + }, + }; + return finalizePacket(packet, originalBytes, parsed.lines.length, retainedLines.length); +} + +function normalizeConfiguration(options = {}) { + const maxOutputBytes = options.maxOutputBytes ?? null; + if (maxOutputBytes != null && (!Number.isInteger(maxOutputBytes) || maxOutputBytes < 1)) { + throw new InputError('maxOutputBytes must be a positive integer or null'); + } + const semantic = { + max_output_bytes: maxOutputBytes, + policy_revision: POLICY_REVISION, + reducer_version: REDUCER_VERSION, + }; + return { + maxOutputBytes, + identity: sha256(Buffer.from(canonicalJson(semantic), 'utf8')), + }; +} + +export function reduceTestRun(input, options = {}) { + const invocation = normalizeInvocation(input); + const configuration = normalizeConfiguration(options); + const parsed = parseTranscript(invocation.streams); + const reasons = []; + evidenceStatus(invocation, parsed, reasons); + reasons.sort(); + + const full = buildPacket({ invocation, parsed, configuration, baseReasons: reasons, mode: 'full' }); + if (configuration.maxOutputBytes == null || serializePacket(full).length <= configuration.maxOutputBytes) { + return full; + } + + const references = buildPacket({ invocation, parsed, configuration, baseReasons: reasons, mode: 'references' }); + if (serializePacket(references).length <= configuration.maxOutputBytes) return references; + + const compact = buildPacket({ invocation, parsed, configuration, baseReasons: reasons, mode: 'compact' }); + if (serializePacket(compact).length <= configuration.maxOutputBytes) return compact; + + throw new PayloadCeilingError({ + disposition: 'NEEDS_RAW_EVIDENCE', + input_hash: streamDigest(invocation.streams), + max_output_bytes: configuration.maxOutputBytes, + minimum_safe_packet_bytes: serializePacket(compact).length, + raw_evidence_ref: invocation.source.rawEvidenceRef, + reason: 'PAYLOAD_CEILING_TOO_SMALL', + }); +} + +export function serializePacket(packet) { + return Buffer.from(`${canonicalJson(packet)}\n`, 'utf8'); +} diff --git a/tests/README.md b/tests/README.md index 1795ee7..a281e51 100644 --- a/tests/README.md +++ b/tests/README.md @@ -1,3 +1,10 @@ # Test strategy -Tests must map to named invariants in SPEC.md, include adversarial failure cases, and separate implementation correctness from evidence for the broader hypothesis. +`reducer.test.js` maps to the named invariants in `SPEC.md`. It covers exact +determinism, large success reduction, single and multiple failures, assertions, +stacks, crashes, timeouts, aggregate correctness, stdout/stderr provenance, ANSI, +Unicode, malformed UTF-8 and text, final-newline boundaries, exact measurements, +payload ceilings, raw escalation, CLI failure behavior, and the complete +synthetic conformance corpus. + +These tests establish prototype behavior, not the EXP-001 hypothesis. diff --git a/tests/reducer.test.js b/tests/reducer.test.js new file mode 100644 index 0000000..6d28f15 --- /dev/null +++ b/tests/reducer.test.js @@ -0,0 +1,309 @@ +import assert from 'node:assert/strict'; +import { spawnSync } from 'node:child_process'; +import test from 'node:test'; +import { fileURLToPath } from 'node:url'; +import { + INPUT_PROTOCOL, + InputError, + PACKET_PROTOCOL, + POLICY_REVISION, + PayloadCeilingError, + REDUCER_VERSION, + canonicalJson, + reduceTestRun, + serializePacket, +} from '../src/reducer.js'; +import { conformanceReport, corpus, executeFixture } from '../fixtures/corpus.js'; + +const cliPath = fileURLToPath(new URL('../bin/context-firewall.js', import.meta.url)); + +function fixture(name) { + const value = corpus.find((item) => item.name === name); + assert.ok(value, `missing fixture ${name}`); + return value; +} + +test('rejects an unsupported input protocol', () => { + assert.throws( + () => reduceTestRun({ protocol_version: 'other', streams: [] }), + (error) => error instanceof InputError && error.code === 'INVALID_INPUT', + ); +}); + +test('same input and configuration produce byte-identical output', () => { + const input = fixture('failure/stack-trace').input; + const first = serializePacket(reduceTestRun(input, { maxOutputBytes: 10_000 })); + const second = serializePacket(reduceTestRun(input, { maxOutputBytes: 10_000 })); + assert.deepEqual(first, second); +}); + +test('canonical output contains no generated time or random identity', () => { + const packet = reduceTestRun(fixture('normal/small-all-pass').input); + const output = serializePacket(packet).toString('utf8'); + assert.equal(/timestamp|created_at|generated_at|random|uuid/i.test(output), false); + assert.equal(packet.operation_id, 'op-synthetic-001'); +}); + +test('packet identifies exact protocol, reducer, policy, and configuration', () => { + const packet = reduceTestRun(fixture('normal/small-all-pass').input); + assert.equal(packet.protocol_version, PACKET_PROTOCOL); + assert.equal(packet.receipt.reducer.version, REDUCER_VERSION); + assert.equal(packet.receipt.configuration.policy_revision, POLICY_REVISION); + assert.match(packet.receipt.configuration.identity, /^sha256:[0-9a-f]{64}$/); +}); + +test('large all-pass output is substantially reduced with correct aggregates', () => { + const packet = reduceTestRun(fixture('normal/large-all-pass').input); + assert.deepEqual(packet.decision_evidence.counts, { passed: 1500, failed: 0, skipped: 0, total: 1500 }); + assert.equal(packet.decision_evidence.status, 'passed'); + assert.ok(packet.receipt.measurements.reduced_bytes < packet.receipt.measurements.original_bytes * 0.1); + assert.equal(packet.receipt.suppressed.categories.successful_test, 1500); +}); + +test('a failed test retains identity, message, assertion, and location', () => { + const packet = reduceTestRun(fixture('failure/one').input); + const [failure] = packet.decision_evidence.failures; + assert.equal(failure.identity, 'adds values'); + assert.equal(failure.header.category, 'failed_test'); + assert.deepEqual(failure.details.map((item) => item.category), [ + 'failure_message', 'assertion', 'assertion', + ]); + assert.equal(failure.details[0].text.includes('expected two'), true); +}); + +test('multiple failures remain independent and ordered', () => { + const packet = reduceTestRun(fixture('failure/several').input); + assert.deepEqual(packet.decision_evidence.failures.map((failure) => failure.identity), [ + 'first', 'second', 'third', + ]); + assert.equal(packet.decision_evidence.counts.failed, 3); +}); + +test('pass, fail, and skipped counts are aggregated correctly', () => { + const mixed = reduceTestRun(fixture('failure/mixed-pass-fail').input); + const skipped = reduceTestRun(fixture('normal/skipped-tests').input); + assert.deepEqual(mixed.decision_evidence.counts, { passed: 2, failed: 1, skipped: 0, total: 3 }); + assert.deepEqual(skipped.decision_evidence.counts, { passed: 1, failed: 0, skipped: 2, total: 3 }); +}); + +test('failure stack lines retain their classified evidence', () => { + const packet = reduceTestRun(fixture('failure/stack-trace').input); + const details = packet.decision_evidence.failures[0].details; + assert.equal(details.some((item) => item.category === 'stack_trace'), true); + assert.equal(details.some((item) => item.text.trim() === '---'), true); + assert.equal(details.some((item) => item.text.trim() === '...'), true); + assert.equal(details.some((item) => item.text.includes('public/test.js:4:5')), true); +}); + +test('stdout and stderr provenance are distinct and hashed together', () => { + const item = fixture('edge/mixed-stdout-stderr'); + const packet = reduceTestRun(item.input); + assert.deepEqual(packet.receipt.source.streams.map((stream) => stream.name), ['stdout', 'stderr']); + assert.match(packet.receipt.input_hash, /^sha256:[0-9a-f]{64}$/); + const changed = structuredClone(item.input); + changed.streams[1].data += 'WARNING: another\n'; + assert.notEqual(reduceTestRun(changed).receipt.input_hash, packet.receipt.input_hash); +}); + +test('ANSI is ignored for classification but retained in evidence bytes', () => { + const packet = reduceTestRun(fixture('edge/ansi').input); + assert.equal(packet.decision_evidence.failures[0].identity, 'colored failure'); + assert.equal(packet.decision_evidence.failures[0].details[0].text.includes('\u001b[31m'), true); +}); + +test('superficial PASS, FAIL, error, and warning words do not become verdicts', () => { + const packet = reduceTestRun(fixture('edge/superficial-pass-fail').input); + assert.equal(packet.decision_evidence.status, 'passed'); + assert.equal(packet.decision_evidence.failures.length, 0); + assert.equal(packet.decision_evidence.warnings.length, 0); +}); + +test('runner crash and timeout evidence are explicit', () => { + const crash = reduceTestRun(fixture('process/runner-crash').input); + const timeout = reduceTestRun(fixture('process/timeout').input); + assert.equal(crash.decision_evidence.fatal_errors.length, 1); + assert.equal(crash.decision_evidence.process.exit_code, 2); + assert.equal(timeout.decision_evidence.timeouts.length, 1); + assert.equal(timeout.decision_evidence.process.exit_code, 124); +}); + +test('unexplained nonzero exit requires raw evidence', () => { + const packet = reduceTestRun(fixture('process/nonzero-unrecognized').input); + assert.equal(packet.decision_evidence.status, 'failed'); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); + assert.ok(packet.decision_evidence.reason_codes.includes('NONZERO_EXIT_WITHOUT_RECOGNIZED_FAILURE')); +}); + +test('interrupted output retains recognized evidence and escalates', () => { + const packet = reduceTestRun(fixture('process/interrupted-truncated').input); + assert.equal(packet.decision_evidence.failures[0].details[0].text.includes('cut off'), true); + assert.ok(packet.decision_evidence.reason_codes.includes('INTERRUPTED_OUTPUT')); + assert.equal(packet.receipt.reduction_complete, false); +}); + +test('malformed and ambiguous text cannot become false certainty', () => { + const packet = reduceTestRun(fixture('edge/malformed-output').input); + assert.equal(packet.decision_evidence.status, 'indeterminate'); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); + assert.equal(packet.receipt.unclassified_evidence_present, true); + assert.equal(packet.decision_evidence.unclassified_evidence.length, 2); +}); + +test('contradictory TAP aggregates require raw evidence', () => { + const input = structuredClone(fixture('normal/small-all-pass').input); + input.streams[0].data = input.streams[0].data.replace('# pass 2', '# pass 1'); + const packet = reduceTestRun(input); + assert.equal(packet.decision_evidence.status, 'indeterminate'); + assert.ok(packet.decision_evidence.reason_codes.includes('AGGREGATE_CONTRADICTION')); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); +}); + +test('invalid UTF-8 is hash-addressed and escalated, never dropped', () => { + const input = structuredClone(fixture('normal/small-all-pass').input); + input.streams = [{ name: 'stdout', encoding: 'base64', data: Buffer.from([0xff, 0xfe]).toString('base64') }]; + const packet = reduceTestRun(input); + assert.equal(packet.decision_evidence.status, 'indeterminate'); + assert.ok(packet.decision_evidence.reason_codes.includes('MALFORMED_UTF8')); + assert.equal(packet.decision_evidence.unclassified_evidence[0].byte_count, undefined); + assert.match(packet.decision_evidence.unclassified_evidence[0].text, /non-UTF-8 2 bytes/); +}); + +test('missing raw reference is distinct from reducer destruction', () => { + const packet = reduceTestRun(fixture('provenance/raw-reference-absent').input); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); + assert.equal(packet.receipt.raw_evidence.escalation_available, false); + assert.equal(packet.receipt.raw_evidence.source_evidence_disposition, 'PRESERVATION_UNCONFIRMED'); + assert.equal(packet.receipt.raw_evidence.destroyed_by_reducer, false); +}); + +test('missing source identity is explicit and indeterminate', () => { + const packet = reduceTestRun(fixture('provenance/source-missing').input); + assert.equal(packet.receipt.source.id, null); + assert.ok(packet.decision_evidence.reason_codes.includes('SOURCE_IDENTITY_MISSING')); + assert.equal(packet.decision_evidence.status, 'indeterminate'); +}); + +test('missing operation and exit identities fail closed', () => { + const input = structuredClone(fixture('normal/small-all-pass').input); + delete input.operation_id; + input.process.exit_code = null; + const packet = reduceTestRun(input); + assert.equal(packet.decision_evidence.status, 'indeterminate'); + assert.ok(packet.decision_evidence.reason_codes.includes('OPERATION_ID_MISSING')); + assert.ok(packet.decision_evidence.reason_codes.includes('EXIT_STATUS_MISSING')); +}); + +test('an exact payload boundary is honored byte for byte', () => { + const result = executeFixture(fixture('payload/exact-boundary')); + assert.equal(result.pass, true); + assert.equal(serializePacket(result.packet).length, result.options.maxOutputBytes); + assert.equal(result.packet.receipt.payload_limit.affected, false); +}); + +test('payload pressure never silently truncates critical evidence', () => { + const item = fixture('payload/critical-too-large'); + const packet = reduceTestRun(item.input, item.options); + const output = serializePacket(packet); + assert.ok(output.length <= item.options.maxOutputBytes); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); + assert.ok(packet.decision_evidence.reason_codes.includes('PAYLOAD_LIMIT_CRITICAL_EVIDENCE_EXCEEDED')); + assert.equal(packet.decision_evidence.failures[0].identity, 'oversized-critical'); + assert.equal('details' in packet.decision_evidence.failures[0], false); + assert.equal(packet.receipt.payload_limit.affected, true); +}); + +test('payload priority replaces warning text with an addressable reference first', () => { + const input = structuredClone(fixture('normal/small-all-pass').input); + input.streams[0].data += `WARNING: ${'bounded-warning'.repeat(600)}\n`; + const packet = reduceTestRun(input, { maxOutputBytes: 3_000 }); + assert.equal(packet.decision_evidence.status, 'passed'); + assert.equal(packet.decision_evidence.disposition, 'NEEDS_RAW_EVIDENCE'); + assert.ok(packet.decision_evidence.reason_codes.includes('PAYLOAD_LIMIT_OMITTED_NONCRITICAL_TEXT')); + assert.equal(packet.decision_evidence.warnings.length, 1); + assert.equal('text' in packet.decision_evidence.warnings[0], false); + assert.ok(packet.decision_evidence.warnings[0].byte_count > 3_000); + assert.ok(serializePacket(packet).length <= 3_000); +}); + +test('a ceiling below the safe packet size fails with machine details and no packet', () => { + const item = fixture('payload/safe-packet-impossible'); + assert.throws( + () => reduceTestRun(item.input, item.options), + (error) => error instanceof PayloadCeilingError + && error.details.disposition === 'NEEDS_RAW_EVIDENCE' + && error.details.minimum_safe_packet_bytes > item.options.maxOutputBytes, + ); +}); + +test('original and reduced byte and event measurements are exact', () => { + const input = fixture('normal/small-all-pass').input; + const packet = reduceTestRun(input); + const output = serializePacket(packet); + const expectedInputBytes = input.streams.reduce((sum, stream) => sum + Buffer.byteLength(stream.data), 0); + assert.equal(packet.receipt.measurements.original_bytes, expectedInputBytes); + assert.equal(packet.receipt.measurements.reduced_bytes, output.length); + assert.equal( + packet.receipt.measurements.retained_evidence_count + packet.receipt.measurements.suppressed_evidence_count, + packet.receipt.measurements.original_event_count, + ); +}); + +test('duration is caller-supplied semantic evidence and no runtime latency is hashed', () => { + const packet = reduceTestRun(fixture('normal/small-all-pass').input); + assert.equal(packet.decision_evidence.process.duration_ms, 12.5); + assert.equal('processing_latency_ms' in packet.receipt.measurements, false); +}); + +test('missing final newline has an exact event count', () => { + const item = fixture('edge/missing-final-newline'); + const packet = reduceTestRun(item.input); + assert.equal(packet.receipt.measurements.original_event_count, item.input.streams[0].data.split('\n').length); +}); + +test('CLI reads JSON from stdin and emits the canonical packet', () => { + const input = fixture('normal/small-all-pass').input; + const result = spawnSync(process.execPath, [cliPath, 'reduce'], { + encoding: 'utf8', + input: JSON.stringify(input), + }); + assert.equal(result.status, 0, result.stderr); + const expected = serializePacket(reduceTestRun(input)).toString('utf8'); + assert.equal(result.stdout, expected); +}); + +test('CLI reports malformed JSON as a machine-readable error', () => { + const result = spawnSync(process.execPath, [cliPath, 'reduce'], { + encoding: 'utf8', + input: '{broken', + }); + assert.equal(result.status, 2); + assert.deepEqual(JSON.parse(result.stderr), { + code: 'INVALID_INPUT', + message: 'input must be valid JSON', + }); + assert.equal(result.stdout, ''); +}); + +test('CLI keeps an impossible-ceiling error off model-visible stdout', () => { + const item = fixture('payload/safe-packet-impossible'); + const result = spawnSync(process.execPath, [cliPath, 'reduce', '--max-bytes', '64'], { + encoding: 'utf8', + input: JSON.stringify(item.input), + }); + assert.equal(result.status, 2); + assert.equal(result.stdout, ''); + const error = JSON.parse(result.stderr); + assert.equal(error.code, 'PAYLOAD_CEILING_TOO_SMALL'); + assert.equal(error.disposition, 'NEEDS_RAW_EVIDENCE'); +}); + +test('all 30 synthetic fixtures conform', () => { + const report = conformanceReport(); + assert.equal(report.fixture_count, 30); + assert.equal(report.conformance, 'PASS'); + assert.equal(report.fixtures.every((item) => item.conformance === 'PASS'), true); +}); + +test('conformance output is deterministic', () => { + assert.equal(canonicalJson(conformanceReport()), canonicalJson(conformanceReport())); +});