Parent roadmap: #740. Governing completed specification: #741 (version 741-v1.0). Complete: the bounded offline prototype was implemented by Codex CLI and accepted on 5 September 2026 at commit 0c7321f0. This is offline acceptance only. The complete normative documents and source hashes are embedded in #741.
Prepared Codex CLI assignment — offline prototype only
Status: completed by Codex CLI 0.153.1 in an isolated external experiment repository. The frozen #741 documents govern the work. This assignment authorizes offline implementation and synthetic checks only; real model experiments and Madar production changes remain outside scope.
Read CONTRACT.md, PROMPTS.md, EVALUATOR-CONTROLS.md, COMPARATOR-SELECTION.md and their frozen manifest. Build only the offline evidence-reconciliation evaluation prototype needed to determine whether the prospective measurement can be executed faithfully. You are not implementing a Madar feature, building a ranker, repairing #739 or claiming an experiment result.
Source and work boundary
Use a dedicated experiment directory outside the Madar product repository and control plane, with its own Git history. The integration-base reference for source inspection is origin/next@72ecb4aa72899c5fa1ba4e2c27795070e74871eb; do not edit it. Do not create another Madar worktree solely for this external prototype. Preserve the root checkout, all existing worktrees and historical evidence.
Allowed outputs: a strict execution-manifest and response loader; a deterministic provenance/ID/link validator; receipt/cost/time accounting and admission checks; a semantic-adjudication input format that never pretends structural checks establish truth; a report generator; inert synthetic fixtures/controls; tests; a reproducible dependency lock and README. Keep these under src/, tests/, fixtures/, dependency manifests and README in the dedicated experiment repository.
Forbidden: Madar production/test changes; archived qualification truth; real target-task construction; live source-provider installation/indexing; calls to paid models or benchmark agents; an automated LLM semantic judge; shell/MCP execution capability in the synthetic runner; product integration, dependency upgrades in Madar, PR #742 edits, merge or release.
Implementation and acceptance
- Implement a single public offline command that consumes a versioned manifest plus synthetic receipts/answers/adjudications and either rejects them with a typed reason or produces all three scorecards. It must not have a live model-execution mode in this assignment.
- Enforce actual input identity, exact evidence-ID disposition, final-link validity, physical source lines, dense schema closure and source boundaries. Explicitly leave semantic truth to authenticated adjudication records; never derive readiness from complete fields.
- Implement the frozen accounting equations, run/task/repository reporting, shared-draft allocation versus actual billing, invalid/failed distinction, token/time/cost admission and all stop conditions. Reject absent meter/rate/reservation identities. Do not invent API options or claim a live budget has been enforced from mocked evidence.
- Create and independently check the controls and mutations in EVALUATOR-CONTROLS.md. Include one passing complete synthetic campaign and failing cases through the same loader/report boundary. Demonstrate that semantic failure can coexist with structural success and is preserved in the final recommendation.
- Deliver one frozen source commit/tree, exact dependency/runtime inventory, one command to run the focused control suite, the observed results, mutation-to-owning-control attribution, source/permission limitations, and a sample report explicitly labelled synthetic. Include a proposed live-execution manifest with missing real identities visibly refused, rather than fabricated.
Cap engineering at one working day and one scope. Stop with the exact unresolved contract or capability mismatch if it cannot fit. Do not open a sequence of architecture issues or weaken the contract to make the sample pass. A later reviewer examines this exact offline candidate; real execution requires its own concrete request and price/permission/identity preflight.
The value of this assignment is a falsifiable measurement tool. Passing synthetic controls establishes that tool's bounded bookkeeping behavior only; it does not establish better agent answers, a valid held-out campaign, product recovery, or a Madar advantage over Serena/Native.
Current execution record
The author used the existing configured model, gpt-6-astra / ultra, and Node 22.22.3. Installed CLI 0.149.0 rejected that model before coding; the compatible app-bundled CLI 0.153.1 completed the implementation. No CLI upgrade or global configuration change was made. The earlier launch is retained as a failed startup, not a completed implementation.
The frozen candidate, observed synthetic controls and bounded source reviews are complete. No live experiment or product-quality verdict is implied.
Offline delivery accepted — 5 September 2026
- Source commit:
0c7321f0ae4db61484a7cb4f7fb25447d60ccd84; tree: 9e13741a99f482766eb55d924930d010500f39f0.
- Evaluator/test/document inventory SHA-256:
bb17acb5cb0ff3f86c15a9baf89b4228b909fd65aeea930ee9a28c17c30556ba.
- Node 22.22.3, no external dependencies. Focused command from the standalone checkout:
node tests/controls.mjs.
- 157/157 public-command assertions across C01–C23; 61/61 authored expectation-consistency checks. Both complete fixed and system synthetic reports produce all three scorecards.
- Two independent AI source reviewers checked all 17 frozen semantic specimens and found no remaining concrete blocker. A separate bounded code review verified the accounting/chronology corrections and final expectation guard. These were accurately labelled AI reviews, not fabricated human attestations.
- The lead repeated the focused command and package verification at this exact commit in a separate checkout: both exited 0, results reproduced, and Git remained clean. The final index covers 1,685 evidence files, with historical failures explicitly separated from current evidence.
- The source, reports, runtime inventory, input/output control attribution, review records, before/after reproductions and full archive are retained locally with this Codex task under
outputs/offline-743/. Canonical delivery record: DELIVERY.md / ACCEPTANCE.json; standalone source: work/743-offline/source.
- Full archive SHA-256:
e494df1ed991d029b6498228e0f59428f9c3fd78f30dca299263d49776ef8e3c.
The repaired checks reconcile raw visible output with billing, include assembly/validation in campaign chronology and reservations, and require completed development adjudication/admission before held-out access. Source-checked fixtures now preserve signed-zero equivalence, precise uncertainty, real oracle-defect semantics and consistent citation decisions. Changed statement bytes/targets cannot silently inherit authored judgments. The frozen #741 documents and thresholds were not changed.
Synthetic samples intentionally retain issue_complete=false and pending genuine-review fields; their deterministic test keys cannot authenticate the real external review records. This issue's closure records the separately reviewed offline deliverable, not a live campaign outcome. The proposed live manifest is refused with missing real identities, attributable billing and enforcement still unverified. Madar's actual quality/speed/cost and product decision remain open in #740. No provider install, Madar production edit, #739 resumption, PR742 edit, merge or release occurred. The root checkout and 38 registered worktrees were preserved during this assignment.
Parent roadmap: #740. Governing completed specification: #741 (version 741-v1.0). Complete: the bounded offline prototype was implemented by Codex CLI and accepted on 5 September 2026 at commit
0c7321f0. This is offline acceptance only. The complete normative documents and source hashes are embedded in #741.Prepared Codex CLI assignment — offline prototype only
Status: completed by Codex CLI 0.153.1 in an isolated external experiment repository. The frozen #741 documents govern the work. This assignment authorizes offline implementation and synthetic checks only; real model experiments and Madar production changes remain outside scope.
Read CONTRACT.md, PROMPTS.md, EVALUATOR-CONTROLS.md, COMPARATOR-SELECTION.md and their frozen manifest. Build only the offline evidence-reconciliation evaluation prototype needed to determine whether the prospective measurement can be executed faithfully. You are not implementing a Madar feature, building a ranker, repairing #739 or claiming an experiment result.
Source and work boundary
Use a dedicated experiment directory outside the Madar product repository and control plane, with its own Git history. The integration-base reference for source inspection is
origin/next@72ecb4aa72899c5fa1ba4e2c27795070e74871eb; do not edit it. Do not create another Madar worktree solely for this external prototype. Preserve the root checkout, all existing worktrees and historical evidence.Allowed outputs: a strict execution-manifest and response loader; a deterministic provenance/ID/link validator; receipt/cost/time accounting and admission checks; a semantic-adjudication input format that never pretends structural checks establish truth; a report generator; inert synthetic fixtures/controls; tests; a reproducible dependency lock and README. Keep these under
src/,tests/,fixtures/, dependency manifests and README in the dedicated experiment repository.Forbidden: Madar production/test changes; archived qualification truth; real target-task construction; live source-provider installation/indexing; calls to paid models or benchmark agents; an automated LLM semantic judge; shell/MCP execution capability in the synthetic runner; product integration, dependency upgrades in Madar, PR #742 edits, merge or release.
Implementation and acceptance
Cap engineering at one working day and one scope. Stop with the exact unresolved contract or capability mismatch if it cannot fit. Do not open a sequence of architecture issues or weaken the contract to make the sample pass. A later reviewer examines this exact offline candidate; real execution requires its own concrete request and price/permission/identity preflight.
The value of this assignment is a falsifiable measurement tool. Passing synthetic controls establishes that tool's bounded bookkeeping behavior only; it does not establish better agent answers, a valid held-out campaign, product recovery, or a Madar advantage over Serena/Native.
Current execution record
The author used the existing configured model,
gpt-6-astra/ultra, and Node 22.22.3. Installed CLI 0.149.0 rejected that model before coding; the compatible app-bundled CLI 0.153.1 completed the implementation. No CLI upgrade or global configuration change was made. The earlier launch is retained as a failed startup, not a completed implementation.The frozen candidate, observed synthetic controls and bounded source reviews are complete. No live experiment or product-quality verdict is implied.
Offline delivery accepted — 5 September 2026
0c7321f0ae4db61484a7cb4f7fb25447d60ccd84; tree:9e13741a99f482766eb55d924930d010500f39f0.bb17acb5cb0ff3f86c15a9baf89b4228b909fd65aeea930ee9a28c17c30556ba.node tests/controls.mjs.outputs/offline-743/. Canonical delivery record:DELIVERY.md/ACCEPTANCE.json; standalone source:work/743-offline/source.e494df1ed991d029b6498228e0f59428f9c3fd78f30dca299263d49776ef8e3c.The repaired checks reconcile raw visible output with billing, include assembly/validation in campaign chronology and reservations, and require completed development adjudication/admission before held-out access. Source-checked fixtures now preserve signed-zero equivalence, precise uncertainty, real oracle-defect semantics and consistent citation decisions. Changed statement bytes/targets cannot silently inherit authored judgments. The frozen #741 documents and thresholds were not changed.
Synthetic samples intentionally retain
issue_complete=falseand pending genuine-review fields; their deterministic test keys cannot authenticate the real external review records. This issue's closure records the separately reviewed offline deliverable, not a live campaign outcome. The proposed live manifest is refused with missing real identities, attributable billing and enforcement still unverified. Madar's actual quality/speed/cost and product decision remain open in #740. No provider install, Madar production edit, #739 resumption, PR742 edit, merge or release occurred. The root checkout and 38 registered worktrees were preserved during this assignment.