Skip to content

roadmap: separate maintenance, evidence use, and Madar product qualification #740

Description

@mohanagy

Current status — 7 Sep 2026

#750 merged (9214c07b); post-merge CI passed 6/6. Allocation repair passes. The complete-evidence goal still fails on discovery and owner-input loss. #740 stays open.

Proof and next diagnosis.

Earlier sections below are historical.


Madar roadmap: reliability, evidence use, and a bounded product decision

Updated 5 September 2026 following the maintainer's request to refine the roadmap, reconcile open issues, clean worktrees, and revisit the code. This issue replaces obsolete planning queues; it preserves the historical decisions and evidence.

Current status — 6 September 2026

Source-preservation candidate remains on HOLD

Local candidate f6d9a83 on next72ecb4aa passes108focused tests, typecheck and build, but the unchanged consumer replay passes4/6: same-file ordering is repaired; rounding and normalization still lose helper source. Fresh independent review returned IMPLEMENTATION-HOLD. The missing branch is an empty lexical-scored symbol list despite available bounded source; the correction stays in the snippet owner. Evidence: #740 (comment)

One author identity was paused/continued and one fresh reviewer completed. This cycle recorded394,864input tokens; aggregate4,491,680 against the conservatively stated4.5million ceiling. Earlier paused-segment usage is explicitly marked observed/incomplete. No additional model dispatch, API-account use, integration, new issue, #734/#739 restart or release. #740 remains open; documentation PR742 stays separate. Earlier comparison results remain historical evidence.

Completed one-task diagnostic — 6 September 2026

The approved Got T3 diagnostic completed all three investigation/review workflows and two fresh blinded semantic graders. Product comparison is inconclusive. The six requirements below belong to this single task; they are not the historical six Madar quality gates.

Configured workflow Grader 1 resolved Grader 2 resolved Internal citation-link errors Investigation + review Recorded input
Madar 5/6 2/6 11 460.427s 271,511
Native 2/6 3/6 15 246.802s 260,305
Serena 2/6 1/6 16 262.880s 207,391

Every investigator used six shell operations and zero MCP calls, despite provider readiness/configuration passing. These times are recorded run-stage sums from one warm, ordered repetition, excluding prior preparation, coordinator gaps and grading. They cannot establish graph-tool quality or a causal speed ranking.

Both graders judged the source explanations largely supported, flagged no critical false claim or semantic citation error within the supplied evidence, and judged every verification plan incomplete. They disagreed materially about requirement coverage; both scores are retained without coordinator adjudication. Separately, deterministic citation-link checks failed for all three answers despite valid JSON/schema and physical citation bounds. Independent data audits reproduced all link errors. No response was repaired or replaced.

Combined predecessor plus diagnostic usage is 3,632,621 input / 123,244 output tokens, within the explicitly approved4million input ceiling. Cached input is a subset, not additional input.19 completed model sessions plus1 retained pre-model aborted slot; coordinator/adviser chat is unmeasured and excluded. All three investigators exceeded the90k input target; source confirmation/both graders also did, and G02 reported8,599 output tokens. These retained v2 resource flags do not constitute a historical gate pass. Actual dollars and subscription-quota attribution remain unknown. Existing ChatGPT-authenticated CLI only; the excluded API account was not used.

Issue740 stays open. This evidence supports separately checking explanation correctness, verification completeness and optional tool adoption before another product-comparison campaign. No new campaign/issue, product patch, issue739 restart, merge or release was launched. PR742 retains its separate disposition. Root changes and38 registered worktrees were preserved. The stopped v1 evidence and earlier v2 snapshot were verified unchanged.

Current local report: outputs/system-diagnostic-740-v2/RESULT.md; structured results: RESULT.json, GRADE-COMPARISON.json, MEASUREMENTS.json; final artifact manifest: EVIDENCE-MANIFEST-FINAL.json. Earlier progress/status sections below are historical snapshots and are superseded by this completed result.

Diagnostic resumed with approved input allowance — 6 September 2026

The maintainer approved the explicit request to raise the combined recorded input ceiling from3million to4million. Carry2,805,241 input tokens forward;1,194,759 remain before further dispatch. Continue the frozen Native and Serena workflows, then the two blinded graders. Task/oracle/prompts/provider configuration, sample/order and the recorded Madar result stay unchanged. Previous pending-status snapshots below are historical. No API account or monetary claim; no sample expansion or historical gate pass. Local BUDGET-DECISION.json and BUDGET-AMENDMENT.md retain the approval and exact scope.

Diagnostic progress: budget decision pending — 6 September 2026

The one-sentence oracle correction and fresh source confirmation are complete. The Madar-enabled investigation and ordinary review finished: six shell reads, zero Madar calls;316.497s investigation and143.929s review;211237+60274 input tokens. Its review passed JSON/schema checks but has11 internal citation-link inconsistencies (eight mismatched evidence/record links and three missing required supporting citations). No answer was repaired. Semantic grading remains pending. These observations do not establish graph-retrieval quality or a three-system ranking.

Combined recorded input is2805241, leaving194759 under the existing3million ceiling. The coordinator asked the maintainer whether to raise the combined ceiling to4million to finish Native, Serena and the two blinded graders. No increase is assumed; those jobs have not started. No recorded model job is still running. The successor remains unfinished and can resume with the same frozen inputs after that budget decision. Original v1 remains stopped and immutable. Local progress: outputs/system-diagnostic-740-v2/PROGRESS.md and PROGRESS.json. No API account, product edit, merge/release or issue739 restart.

One-task subscription diagnostic authorized — 6 September 2026

The maintainer proceeded with one expected-answer correction followed by a small three-system comparison. Codex CLI corrected only T3-O2.requirement in a new oracle version; an independent fresh source reviewer admitted the corrected T3 pair after checking all six obligations. All other fields and the original stopped v1 are unchanged.

The successor740-diagnostic-v2 uses the same Got T3 task, one repetition and three investigation/review workflows, in the original T3 order: Madar, Native, Serena. Prompts, oracle, source, runtime/configuration and order are frozen before experimental answers. All use the existing ChatGPT-authenticated recorder and identical ordinary review A. No owner API account/key, billing/count endpoint, new implementation or product change. This small diagnostic cannot satisfy the historical six-task gates or establish an overall winner. Per-session90k/8k overruns are disclosed resource flags in this exploratory successor rather than discarded answers; the combined3million-input/600k-output ceiling and prior usage remain in force.

Previous and correction/check sessions have consumed2,533,730 input tokens;466,270 remain before further dispatch. Actual dollars, exact subscription quota and coordinator/adviser chat usage remain unmeasured. Two blinded graders are planned; unknown usage or exhaustion stops further dispatch. Existing indexes are reused and prior cold setup is reported separately; a one-repetition warm run is not a cold-speed ranking. Local protocol and freeze: outputs/system-diagnostic-740-v2/PROTOCOL.md and EXPERIMENT-FREEZE.json. No merge, release, issue739 restart or new issue is authorized by this diagnostic.

Subscription comparison stopped before scoring — 6 September 2026

Protocol 740-agent-system-v1 stopped at task admission. Zero scored system workflows ran; no Native/Madar/Serena quality, speed or cost result exists. The third fresh source reviewer rejected a critical expected-answer obligation in Got task T3. Explicit UTF-8 spellings and undefined encoding use TextDecoder; the oracle misstated the decoding rule. A correct system answer could therefore have been penalized. This is a benchmark truth defect, not a Madar product failure. Reviewer references: response.ts and core/index.ts.

T1/T2/T4 were admitted; T3 was rejected; T5/T6 were not source-reviewed after the stop. All original task/oracle versions are retained without repair or replacement. Nine ChatGPT-authenticated CLI preparation sessions completed, using 2,375,165 input tokens (including 1,982,464 cached) and 67,511 output. That consumed 79.17% of the 3 million input allowance before scored work. The coordinator underestimated preparation consumption. The entire planned campaign cannot be claimed feasible within the remainder from this evidence. One additional orchestration attempt stopped before authentication/model dispatch; its slot and error remain recorded. A completed receipt was recovered after a separate coordinator bookkeeping error, with no repeated model call. Actual dollars, precise subscription quota and coordinator/adviser chat usage remain unmeasured.

Madar 0.32.1 and pinned Serena runtimes are installed; static Madar indexes were generated. All nine completed recorder artifact sets, terminal usage and unchanged source manifests verified. Original recorded process IDs are no longer live. No owner API account, billing page, API fallback, Madar product edit, worktree deletion, merge, release or issue739 restart occurred.

Current disposition: preserve this stopped attempt and the existing recorder. The smallest successor is a versioned correction of the one expected-answer defect and a small bounded diagnostic on admitted tasks, with explicit limits before dispatch. No successor is running and no full-campaign pass is implied. No new framework or issue is needed. Issue740 remains open; 741/743 remain complete; 739 remains blocked. Local result and stop receipt: outputs/system-comparison-740-v1/RESULT.md and STOP-RECEIPT.json in the coordinating task. The full original protocol remains below, SHA256 a1f43833ad577e03b351669623d1517203ea5745f0893ff36f9cbe0b0ecb8cb2.

The approved741-v1.1 offline API adapter is complete and accepted for that bounded scope at source 4c270e12d9c1c4541244c443218058580f9ae73c, tree 963a6f8a2e672035bd187b4cd2a19b0536f7ea9c. Codex CLI authored it. The complete stable-source public suite passed93/93 top-level controls, including47/47 mocked HTTP controls; independent staged-input and checkpoint reviews found no remaining material blocker within their stated scopes. This is not a real-account admission or an experimental product result.

Current direction: use existing agents; the owner API account is excluded from Madar testing. API account admission and project selection are withdrawn. The accepted offline adapter is retained and parked. The agent execution capability record and synthetic integration smoke are now delivered below. The next deliverable is a prospective Native/Madar/Serena comparison amendment: identical model/settings, fresh isolated sessions, common grading, complete tool/time/verified-usage recording, and no API-key fallback. Exact actual-dollar attribution and hard pre-request bounds must not be claimed where the agent route does not establish them. Existing quality thresholds and historical results remain unchanged. That direction change alone did not launch fresh curation or implementation; the subsequent maintainer proceed and active protocol above govern current preparation. #741/#743 remain closed, #739 remains blocked, and this umbrella remains open.

Subscription-agent runner delivered — 6 September 2026

The maintainer-authorized recorder/integration assignment is complete. Codex CLI authored source f636c000ab06449d5b971358c5db826a0c7adbac, tree 3b13060593c66ecb01113668ee9624144fe656f1 (19 files). Final synthetic suite20/20 passed. Complete Git bundle and all smoke receipt artifact hashes verified. Full local delivery: outputs/agent-runner-740/DELIVERY.md and DELIVERY.json in the coordinating task.

Live execution used the existing ChatGPT-authenticated CLI0.149.0, requesting gpt-5.6-sol/medium/default. Native ordinary inspection, Madar stable0.32.1 core calls, and pinned Serena symbol lookup all executed on an unrelated two-file synthetic fixture. Each retained raw events, final answer, reported tokens, observer elapsed time and unchanged source manifests. No OpenAI API key/account, count endpoint, billing setup or API fallback was used.

The first Madar attempt is retained as an integration failure: its MCP calls required approval and were refused. CLI added explicit approval for only the existing seven Madar tools; one corrected smoke executed both calls. Retrieve nevertheless returned no matching evidence, so the final answer relied on ordinary source inspection. Connectivity is demonstrated; retrieval usefulness is not. Serena setup separately preserved a cross-file reference miss on the tiny no-tsconfig fixture, while symbol body lookup and same-file references worked. These are integration observations, not a comparative product score.

All three completed paths identified the synthetic arithmetic bug. No timing/token ranking is claimed: prompts explicitly exercised each provider, setup/cache conditions differ and the fixture is trivial. All four smoke attempts are retained; author/correction usage is recorded separately as one-time engineering. Actual dollars, subscription quota use, immutable returned model identity and hard in-flight token/spend bounds remain unestablished. No old cost/quality gate is marked passed.

Madar root status and all38 registered worktrees match the start snapshot. No Madar product edit, worktree removal, fresh benchmark selection, scored campaign, merge, release or #739 resumption occurred. Native/Serena launch/config/input hashes remained unchanged after the Madar-only correction, so their earlier-head smoke observations are labelled and retained without reruns.

The next deliverable is the prospective subscription-based comparison protocol and fresh task custody. It must explicitly resolve the earlier fixed-evidence prerequisite and account for available time/token measurements, unavailable actual-dollar/hard-reservation evidence, and CLI/provider observation limits before scoring. The agent route is executable; the Madar product hypothesis remains untested. No new issue is needed. #740 stays open, #741/#743 stay complete, #739 stays blocked, and the API adapter remains parked.

Intended outcome and present limit

Madar should help a coding agent understand relevant TypeScript/Node behavior, propose and implement correct changes, and reduce the total work required. Local graph accuracy, compact output, and fewer ordinary shell calls are supporting measurements. They do not establish that outcome by themselves.

The accepted #736 result remains PROTOTYPE-KILL-736: two of six read-only investigations met a composite evidence, planning, and benefit rubric. This was not a patch-resolution rate or universal accuracy score. Native also missed critical evidence. The postmortem distinguishes evidence unavailable to the provider from evidence returned and then omitted in the final plan. The R3 latency outlier remains material even though the historical median-efficiency gate passed.

Current queue

Work Disposition Next deliverable
This roadmap and workspace cleanup Coordination Current issue/project/document references, source audit, retained evidence and a recoverable cleanup record.
#741 evidence-use research contract Complete: contract 741-v1.0 Fixed draft/evidence comparison, separate scoring, finite budget and source-checked comparator specified. No experimental result.
#743 offline evaluator prototype Complete: offline candidate 0c7321f0 accepted One external prototype with synthetic controls and accounting; no real model campaign or Madar production change.
741-v1.1 API adapter under #740 Offline implementation complete; execution parked Owner API account excluded from testing; no account-admission or project-selection dependency. Prepare an existing-agent comparison with explicit measurement limits.
#710 test timing and worker-start reliability Independent maintenance Separate deterministic timing behavior from worker lifecycle failures; define a bounded CLI assignment on the chosen integration base.
#697 workspace Git process policy Independent maintenance Refresh the call-site inventory and preserve distinct not-a-repository, command-error, and bound-exceeded outcomes.
#739 static namespace bracket calls Blocked; mechanism stopped Preserve the unresolved baseline defect and both rejected candidates. No further author run or acceptance rerun without a new mechanism decision and corrected, versioned control specification.
Dependency PRs #711, #712, #713, #727, #728 Separate maintenance Exact-head review and cause-specific CI diagnosis before any merge. No dependency upgrade is assumed to fix product quality.

#710 and #697 do not block preparation of the research contract. #710 becomes an execution prerequisite only when a required run cannot be trusted. #739 is not a prerequisite for testing evidence use with an unchanged provider. Do not turn all open issues into one mandatory chain.

Code disposition

  • Retain local source indexing, exact source references, explicit uncertainty, bounded context selection, and existing compiler-backed SPI components.
  • Preserve default-auto topology ownership while documenting it accurately: auto currently retains legacy relations and supplemental SPI metadata. Switching to SPI topology is a separate compatibility decision.
  • Keep [P0 Prototype] Falsify interactive TypeScript evidence navigation against Native exploration #736's TypeScript Language Service navigator as a frozen prototype/control, not a shipped replacement. It has a separate session-freshness limitation that must be addressed before use across live edits.
  • Treat context-pack coverage and recovery as retrieval diagnostics. They do not inspect or validate the consuming agent's final implementation plan.
  • Retain fix: preserve static namespace bracket calls in default-auto call graphs #739's candidates as diagnostic evidence. The r2 heritage traversal regression, invalid decorator positive control, and incomplete validation are distinct findings.
  • Defer broad ranker rewrites, deeper graph expansion, a new general memory product, hosted services, language expansion, and automatic promotion of unreleased next to main.

Source basis: auto ownership, SPI compiler setup, pack bookkeeping, recovery boundary.

Prospective evaluation rules

Report three separate results: (1) tool/extraction correctness, (2) final-answer and plan correctness, and (3) complete-task benefit. Reserve executable patch/test success for experiments that actually produce and verify patches.

Freeze task scope, independently reviewed expected behaviors, admissible semantic equivalents, scoring procedures, model/configuration, compared arms, execution-validity rules, complete cost accounting, per-task regression bounds, and a finite run budget before accessing fresh held-out tasks. Trace available source -> retrieved evidence -> returned evidence -> final use. A truthful unknown is not found evidence; harmless naming differences are not automatically semantic failures.

Measure all ordinary and MCP operations, startup/index/refresh time, complete elapsed time, model input/cache/output accounting, and actual or explicitly estimated charges. Report individual outliers as well as aggregates. A reduction in one call category must not conceal a severe elapsed-time or cost regression.

Include Native and one relevant existing tool in any later comparative experiment. Do not compare unrelated vendor percentages with the old 2/6 count. Keep the exposed #736 tasks diagnostic; they cannot become fresh holdouts or tuning targets.

The evidence-reconciliation mechanism hypothesis remains unproved and its API execution route is parked. The proposed next experiment is the existing agent-based Native/Madar/Serena system comparison. A prospective amendment must explicitly change sequencing, freeze the same neutral review across systems, and define supported resource measurements before fresh tasks or dispatch. Keep this product comparison separate from the unrun fixed-evidence mechanism test; do not imply that either has passed.

Research informing the decision

  • CodePlan, FSE 2024: dependency-aware planning and propagation of related edits.
  • RepoGraph, ICLR 2025: relevant graph neighborhoods can help; more context and summarization can also reduce correctness.
  • SWE-agent, NeurIPS 2024: action granularity and result presentation change agent behavior and task outcomes.
  • Agentless: separate localization, repair, and validation; generated reproduction tests also need validation against correct behavior.
  • Lost in the Middle, TACL 2024: evidence presence and evidence use are different measurements.

These establish mechanisms worth assessing, not a guarantee that a new Madar architecture will succeed.

History and authority

Keep #734 closed as rejected/not planned; #735 and #736 closed with their accepted terminal decisions; and #738 closed as a completed narrow feasibility investigation. Closed does not mean every underlying defect was fixed. No old benchmark is rescored and no rejected candidate is promoted by this roadmap.

The lead owns planning, GitHub/control-plane reconciliation, documentation, and cleanup. Codex CLI owns any subsequently scoped production implementation and tests. A planning issue is not permission to launch a benchmark, resume #739, merge, release, or change package tags. No release number or date is promised.

Completion of this umbrella requires an explicit product decision supported by the new contract, plus disposition of any resulting work. Merely opening issues or writing this document does not establish product recovery.

Roadmap reconciliation delivered — 5 September 2026

The project board now contains the five current work items and retains all 69 historical cards. Existing issues #697, #710 and #739 remain open with current dispositions; their original bodies and comments are preserved. Four obsolete milestones with no open issues were retired explicitly as planning buckets, not shipped releases.

Draft PR #742 proposes the contributor-roadmap update against next, retaining the old text as a historical archive. The existing focused documentation test passed on Node 22.22.3. This is documentation-only; full PR CI and integration are tracked on the draft.

All 39 original worktrees were inventoried. Two redundant checkouts were removed after complete recovery archives, hash verification and fresh process checks; one isolated documentation worktree was added, leaving 38. Three backup refs and a verified Git bundle retain previously unreferenced history. Unique work, evaluation evidence, referenced performance inputs and active process worktrees remain retained with recorded conditions. The user root checkout was preserved.

The source review and local cleanup/verification records are retained with this task. These coordination results do not resolve #739 or establish product recovery. The #741 specification is complete. The bounded offline Codex CLI prototype in #743 is complete; the product decision remains open.

Specification completed — 5 September 2026

#741 now contains the full frozen specification and document hashes. The decision is SIMPLIFY: test downstream evidence reconciliation against an ordinary review using identical evidence and a common draft, before any provider-specific claim. The later conditional provider comparison uses the same review protocol for Native, stable Madar core and pinned Serena. Quality/source correctness, semantic plan correctness and complete cost/time remain separate.

#743 is the one prepared follow-on assignment: build and verify an offline evaluator prototype from synthetic fixtures within one working day. The maintainer-authorized implementation is complete. Two AI source reviews, a separate code review and a fresh-checkout rerun support offline acceptance at 0c7321f0; 157 public-command assertions and 61 expectation-consistency checks pass. Contract completion neither resumes #739 nor authorizes a model campaign.

Draft PR #742 has completed CI with five passing lanes; Ubuntu Node22 failed the dependency security audit for fast-uri. The failure is recorded separately from #710, with no rerun or dependency change from this research task. The draft roadmap queue must be refreshed to these issue dispositions before eventual integration; its historical snapshot is not a live status board.

Offline evaluator completed — 5 September 2026

#743 is complete for its bounded offline scope at source commit 0c7321f0ae4db61484a7cb4f7fb25447d60ccd84. Its issue body records exact source/tree/inventory identities, synthetic results, two independent AI source reviews, code review, repeatability evidence and local artifact retention. The seventeen source fixtures are checked; no real comparative campaign has run. Closure establishes the prototype's bounded bookkeeping behavior, not Madar recovery or an advantage over Native/Serena.

The next product-research decision concerns whether a concretely requested real run can satisfy the frozen identity, price and enforcement preflight. No new campaign or author task starts automatically. #710/#697 remain independent maintenance choices, #739 remains blocked, and PR742 retains its separate draft/CI disposition.

Historical preflight before amendment approval — 6 September 2026

STOP dispatch under frozen 741-v1.0. CLI0.149.0 and Node22.22.3 are installed and hashed. Detailed token meters exist, but the matching CLI source accounts for completed usage before checking its rollout budget. That mechanism does not establish the contract's required worst-case reservation before dispatch. Actual campaign charge attribution and a complete live execution manifest also remain unverified. This is an execution-admission finding; no new Madar experiment has failed.

Source: 0.149 budget accounting. The installed native executable is SHA256 f4a74117b8142cda581c95ff753abf4508b5636d89682c1ed77e4a9249af8963. Generated installed protocol and pinned-source inspection are retained locally. The five frozen specification hashes and accepted offline source remain unchanged.

At this historical checkpoint, the lead recommended one prospective consumer amendment, then unapproved: use a narrowly scoped Responses API client for the fixed-evidence experiment, with Codex CLI still writing the implementation. Official documentation provides request-aware input counting, an output cap including reasoning, and a per-run reservation pattern. API cache-write charges require separate actual-cost accounting; the original reference index must not be represented as the actual bill.

At that checkpoint, a draft amendment and one bounded offline CLI assignment had received a bounded advisory consistency review. They were not yet effective or dispatched; the subsequent approval and delivery below supersede that planning status. They preserve quality gates and make no claim about Codex CLI/Madar/Serena end-to-end performance. Any paid execution requires a later concrete, supported launch and explicit charge authorization. Runtime capability should have been checked before commissioning #743; its accepted synthetic result nevertheless remains valid for its completed scope.

No new issue was opened. #741/#743 remain closed, #739 remains blocked, and #740 remains open. No fresh tasks, benchmark requests, count probes, provider installations, production edits or worktree removals occurred. Local decision and drafts: outputs/preflight-741-2026-09-06/ in the coordinating Codex task; no credentials are included.

API consumer amendment approved — 6 September 2026

The maintainer approved741-v1.1 and one bounded offline Codex CLI assignment. The original741-v1.0 documents,743 accepted candidate and preflight decision remain preserved. This approval changes the fixed-evidence experiment consumer to a small Responses API client; Codex CLI remains the implementation author. It does not authorize generation/count probes, paid execution, fresh task access or the conditional provider comparison.

The standalone implementation begins from743 source 0c7321f0ae4db61484a7cb4f7fb25447d60ccd84. It is outside the Madar product repository and is tracked here without a new issue. Deliver one adapter, durable budget controller and public-entry mocked controls within one working day. Existing quality gates and session/global limits remain; API actual-cost accounting includes separately billed cache writes and applicable additional charges. The real manifest must remain refused until its supported billing/identity/authorization prerequisites exist.

Approved amendment SHA256: 905396479ada6e3dd301d2cc7b6692a5c2a95dcaf91062b94526a5ab5c146f7a.
Approved CLI assignment SHA256: acfa9b186bb5ceac9ab19ec1df2fa591bb8de3e97d53b8a47067cf2622bf4e7b.
Local authoritative copies: outputs/api-adapter-741/contract/AMENDMENT.md and CLI-ASSIGNMENT.md, with approval/proposal hashes in APPROVAL.json.

The exact approved amendment follows for reviewability:

Approved 741-v1.1 amendment — fixed-evidence consumer only

Status: APPROVED AND EFFECTIVE for amendment preparation and bounded offline implementation, by the maintainer’s “approved” reply on6September2026. Preserve the original five #741 documents and accepted #743 source. No outcome has been measured and no fresh task has been selected through this preflight.

Purpose: make the fixed-evidence mechanism experiment executable with documented request bounds. Codex CLI remains the implementation author. This proposal changes the evaluated consumer from Codex CLI0.149.0 to a small Responses API client. It makes no automatic amendment to the later provider comparison.

Explicit changes

  1. Use exact model string gpt-5.6-sol, reasoning medium, API service_tier=default, synchronous non-streaming Responses calls, no automatic retries, no background execution, no previous-response/conversation linkage and no model substitution. Current documentation lists this model string as its snapshot; no distinct dated immutable identifier has been established. Record and compare every returned model/tier, and retain the contract's alias-drift disclosure.
  2. Freeze the API client source/tree, exact dependency lock and runtime, request serialization, API response schema, permissions and pricing identity before fresh tasks. Node22.22.3 remains the proposed runner runtime. Do not silently use the newer desktop CLI as the experiment consumer.
  3. Each D/A/B session becomes exactly one generation request. Render the original common instruction, arm instruction and response schema into the instructions field, preserving their wording and role-specific visibility. Render TASK, EVIDENCE_PACKET and COMMON_DRAFT when applicable as labelled input data in the frozen order. Publish and hash the exact rendering before use. Do not add code-agent system instructions, goals, skill catalogs, memories or continuation prompts. Omit tool definitions and any tool-enabled execution path. Request store=false; do not imply that this is a zero-retention guarantee.
  4. Count the exact supported model/instructions/input and any other token-bearing fields before generation through the provider's input-count endpoint. Record the count call, its payload identity, duration and any applicable charge. Validate token-field parity between count and generation. No bytes/4 estimate may authorize dispatch. Use the schema text already required by the prompt; if structured-output request fields are added, freeze and count them explicitly first.
  5. Set max_output_tokens=8000, covering all billed output including reasoning. Reject input above90,000. Reserve input, output, session count and both reference/actual cost balances atomically before dispatch. All active reservations consume global allowances. The reservation must survive interruption/restart; uncertain requests are not refunded or retried. Prevent concurrent processes from sharing an unchecked budget. A timeout is not cancellation evidence.
  6. Preserve the frozen normalized reference formula, explicitly labelled a reference index. Calculate API model-token charges separately with four disjoint input/output categories: ordinary=input-cached-write; actual token cost=(4×ordinary+0.40×cached+5×write+20×output)/1,000,000 at the verified short-context/default prices. Reasoning is included once through output. Require presence and valid semantics of cache/write/reasoning fields; missing is not zero. Reserve all input at the highest applicable input rate, presently$5/M. Record actual billing basis, applicable fees/taxes and any count-endpoint or auxiliary charges separately; retain the$30 attributable-actual ceiling. Stop before a paid request if any applicable cost lacks a defensible maximum. Do not equate a usage-derived model-token subtotal with a final invoice.
  7. Freeze one supported caching policy consistently across arms before curation. Default proposal: the exact account/API-supported default policy, with no outcome-dependent cache manipulation and worst-case cache-write reservation. Record any caching request fields explicitly. A different policy requires a documented pre-run choice; never infer zero writes because a meter is absent. Inputs below90,000 avoid the documented>272k long-context price tier.
  8. Preserve failure scoring: incomplete, timeout, refusal, malformed output and other model failures remain failures, not free retries. A known incomplete response with complete usage can settle its charge while retaining its failed score; uncertain charge or receipt integrity stops further dispatch with the reservation retained. Do not copy the cookbook's campaign-stop policy mechanically where it differs from the experiment's failure rules.

Retained experimental rules

The D/A/B comparison, common draft allocation, prompt wording, final-answer schema, fresh task custody, source truth, two development rounds, one permitted B-instruction revision, held-out admission, eight evaluation tasks/three repetitions, blinded semantic scoring and all quality/regression gates stay as frozen. No score is rescored and no threshold is reduced. Keep96 planned sessions and the120 total ceiling, including auxiliary generation work and the existing limited infrastructure replacement. All original per-session/global token/time/cost ceilings remain. Input-count requests are operational requests, not extra generation sessions; their time and any charges are still recorded. No paid canary is part of this amendment or the proposed offline assignment.

Count-call elapsed time belongs to per-session orchestration overhead and full arm cost/time, including common-D allocation. Keep operational receipts distinct from the scored model's zero tool calls. Campaign accounting includes aborted attempts and all campaign auxiliary model work. Preflight/adaptor engineering is disclosed separately as one-time research preparation, not experimental measurements.

Admission before any real request

The maintainer has authorized this versioned amendment and the separately described bounded offline CLI assignment only. A real run still needs its concrete launch manifest, account-specific pricing/charge bound, exact client/runtime identities, full receipt controls and a named authorization covering up to$30 attributable charges. Authentication setup is private and user-controlled; no credential is requested in chat or stored in artifacts. No API call, including a count probe, is authorized by this draft.

Fresh task owners must satisfy the original exclusion rules. The current lead/advisers are not eligible. Avoid spending on curation while execution admission is unresolved. If the adapter or billing route cannot meet the retained requirements, return one precise capability mismatch and stop; do not patch the product or initiate a chain of architecture issues.

Source basis

Responses spending controller supplies the reserve/settle pattern and limitations. Token counting documents request-aware input measurement. Reasoning describes the billed-output cap. Sol and pricing provide model/rate identities. Provider-side spend limits are additional containment and cannot replace per-request reservations because enforcement may be delayed: spend limits.

The approved implementation was delivered by Codex CLI and completed bounded offline review as recorded below. This approval section remains the historical authority for the work. #741/#743 remain closed for their completed scopes; #739 remains blocked.

Offline adapter delivery accepted — 6 September 2026

Source 4c270e12d9c1c4541244c443218058580f9ae73c; tree 963a6f8a2e672035bd187b4cd2a19b0536f7ea9c;46-file inventory SHA256 6b4599acf2576039d761ac5675b3926c0d3676eea69e1bc55ba9e112c5da1444; executable client-source identity f66a4a0d3862f4b7270752311fcc315cdf1196d61decda8ce308c27f715286e3. Runtime Node22.22.3, no third-party dependencies. Source and all36candidate evidence artifacts were byte/hash verified; complete standalone Git bundle verification passed (SHA256 42e029a7dd361c43372f6c616edc5f03a1f1461e3df4fa39829cf07d3d66cf91). Full local delivery: outputs/api-adapter-741/DELIVERY.md and DELIVERY.json in the coordinating Codex task.

The adapter implements exact request/count rendering, complete pre-request reservations, raw meter receipts, separate actual/reference accounting, exclusive durable campaign ownership, failed/uncertain-session retention, and authenticated staged scoring. Candidate1's receipt, timing and strict-ceiling findings were corrected. Candidate2's independent retrospective-judgment finding was corrected with checkpoint v2 and a signed consumed-input seal before subsequent evidence or execution; consumed judgments, B decision, scores and budget cannot be replaced by replay. Both rejected candidates and their evidence remain retained. Original thresholds, prompts, accepted evaluator and contract hashes are unchanged.

Evidence:

  • Author final stable-source suite:93/93 top-level controls, including one aggregate for47/47 injected-HTTP controls, exit0; focused staged-freeze7/7. These are synthetic observations, not model-quality measurements.
  • Independent staged review: original historical replacement plus8related public refusals stopped before loaders/private callbacks/HTTP and retained exact historical bytes;1valid continuation verified the16-record seal before the heldout loader. A later reviewer-helper property error is disclosed; the whole script is not called passing. Lead supplemental row/admission checks completed5/5 separately.
  • Independent checkpoint review:4/4public groups, exit0, covering stale/concurrent continuation, sealed-loader failure, and input-seal/checkpoint signing failures with preserved consumed budget and refused restarts. Its initial reviewer structural-comparison error and corrected attempt are preserved.
  • Independent supplied-real-manifest command: expected exit2 REAL_MANIFEST_SCHEMA, zero credential lookup/count/generation/HTTP/network. The guard remains closed pending real account evidence.

No real API/count request, account credential use, fresh task selection, Madar source change, registered worktree removal, merge or release occurred in this assignment. Madar root HEAD/status and all38registered worktrees match the pre-assignment snapshot; accepted #743 and the five original #741 documents remain unchanged. Separate draft PR742 remains open at5af65b1f with5/6matrix checks passing; no CI retry or merge was launched. This completes one offline engineering deliverable within the approved working-day bound. The product hypothesis and later provider comparison remain untested.

Frozen prospective protocol740-agent-system-v1

Madar subscription system comparison v1

Status: prospective protocol under #740, before fresh task answers or scored runs. The maintainer's proceed after successful subscription integration authorizes carrying the existing-agent comparison forward. This protocol records the changed execution basis explicitly; it does not amend historical results or claim that #741-v1.0/v1.1 passed. Coordinator handles experiment data, source setup, GitHub and verification. Codex CLI remains the implementation author; no new implementation is commissioned by this protocol.

Decision and sequence

Run the direct Native / Madar stable core / Serena comparison now, using ordinary review consistently. The fixed-evidence reconciliation experiment remains unrun and parked. Its success prerequisite is removed for THIS separate system comparison. Do not import its B ledger, two development rounds, >=0.10 mechanism gain, or common-D allocation into these system results. The product question is whether Madar improves source-backed read-only investigation and change planning sufficiently to justify further investment. Patch success, production recovery and release qualification are outside the claim.

Use six fresh tasks from three public TypeScript/Node repositories, two tasks each, two repetitions:36 workflows, each with an investigation and one fresh ordinary review =72 planned experimental sessions. Freeze all task/ground-truth identities, prompts, provider/runtime/configuration and order before the first experimental investigation. No task selection depends on Native/Madar/Serena answers. No tuning or task replacement after outcome access.

Task custody and selection

EXCLUSION-REGISTER.json excludes old qualification, exposed #736 repositories, user projects, forks/mirrors and synthetic fixtures. The coordinator, prior recorder author and existing advisers are ineligible as fresh task owners. A fresh CLI curator receives only this standalone method, exclusion identifiers and metadata of candidate repositories. It freezes the eligible pool and seed ordering before source investigation. Public metadata filtering and source pin retrieval do not inspect task answers.

Seed: madar-740-agent-system-v1-2026-09-06. Rank eligible repository full names by SHA256(seed + newline + lowercase full name). Take the first three, after checking distinct repositories, permissive licenses, usable immutable snapshots and the declared metadata criteria. Preserve every eligibility rejection. Do not search a new pool because a provider performs poorly. If three eligible repositories cannot be established, stop before task curation and publish the missing setup condition.

Allocate task types by selected repository rank: R1 behavior explanation + change-impact plan; R2 behavior explanation + correctness-boundary investigation; R3 change-impact plan + explicit uncertainty investigation. Each asks for current source-backed behavior, a bounded proposed change where applicable, and verification. The uncertainty task must explicitly require a fact unavailable from allowed repository source, without disguising it as a retrievable obligation. All tasks have4–8 meaningful source/task obligations, at least two repository-answerable obligations and at least one critical obligation. Repository-unanswerable requirements form a separate U set.

One fresh curator per repository enumerates public entry points and ranks candidates using SHA256(seed + newline + relative source path + newline + symbol). Inspect at most the first10 ranked candidates to select the first eligible task in each assigned category; retain rejected candidates and source-based reasons. Eligibility means an authentic multi-location behavior or bounded change whose obligations can be independently established, not a manufactured defect or a preferred final wording. Source evidence determines accepted semantic alternatives and justified scope exclusions.

A separate fresh truth checker independently derives obligations, source spans/hashes, criticality, answerability, correct alternatives and required verification from the task and source. A third fresh source reviewer checks the frozen task/oracle pair before admission. A defective/ambiguous specimen is rejected before any measured agent answer; do not soften criticality to admit it. If six eligible tasks cannot be produced within these bounds, stop without a product verdict.

Private task/truth material stays outside investigator source/prompt directories. Root handles identifiers/hashes and dispatch but does not inspect new oracle content or grade responses. Withholding is by explicit role-specific delivery, separate fresh sessions and observed access checks; these same-user CLI processes are not claimed to be protected by an adversarial filesystem sandbox. A model cannot be proved unfamiliar from training. Publish role/readership identities and any exposure, and stop on known leakage. Curators and graders receive no historical postmortem, desired answer or desired winner.

Runtime and observation boundary

Use recorder f636c000ab06449d5b971358c5db826a0c7adbac/tree3b13060593c66ecb01113668ee9624144fe656f1, Node22.22.3 and pinned Codex CLI0.149.0, requesting gpt-5.6-sol, medium reasoning and default/non-Fast processing. Record runtime hashes and requested identity; no immutable returned model identity appeared in smoke. Do not describe the alias as fixed weights. Unexpected reported model/runtime drift stops the campaign; no substitution.

Use existing ChatGPT login only, sanitized environment, ignored user config/rules, fresh ephemeral sessions, disabled memory/skills/hooks/apps/plugins/web/subagents, read-only source, explicit provider configuration. No API key/account, billing page, API count call, API canary, account creation or API fallback. Only the declared repository snapshot may be read. Source instructions are data. Every known source change, outside-source acquisition, undeclared tool use or known task leakage is retained and fails protocol validity.

Native ordinary search/read remains available to every investigation. Provider use is neutral: use the enabled provider when relevant, verify evidence and fall back when needed; count all of it. Do not force successful calls as in the smoke. Madar is published0.32.1, source06b373a447acfce895412ac10eb4e5228c5df0b7, core profile, actual installed TypeScript6.0.3, no optional embeddings/rerank/auto-refresh/experimental tuning. Its admissible tools are retrieve, impact, call_chain, community_overview, pr_impact, graph_stats, graph_summary. Serena is source13ac8c5b1d51873bd148aea440dcb22f85d3a439, local LSP, external per-project state, no memories/activation/editing, exact seven tools from its retained configuration. Record active TypeScript5.9.3 and LSP5.1.3 separately from project source settings and note that compiler versions differ by system.

The CLI also exposes generic MCP resource handlers when a server is configured. No narrow disable switch was established. The stronger old pre-use blocking premise is explicitly replaced here by observable rejection: any generic resource listing/template/read is recorded as a counted MCP operation and fails the execution protocol; it is not an allowed eighth tool. No exact-seven model-menu or prevention-before-access claim is made. Pinned source exposes resource calls in exec JSON; normal model-facing prompt/completion handlers were not found. Preserve initialization instructions and advertised schemas as setup overhead. Unknown event shapes or missing evidence stop further dispatch. This narrow observation policy requires no new client or provider patch.

Workflow and ordinary review

The investigation instruction and plain-text draft follow PROMPTS.md's conditional-system investigator, adding only this protocol's explicit source/tool limits and current task. A fresh review receives TASK, the full returned-observation packet and that workflow's unchanged draft. All providers use the exact same ordinary review A instruction and final JSON semantics from PROMPTS.md; reconciliation is null. Review has zero operations and no MCP/shell acquisition. No review sees hidden truth or another arm's answer.

Assign E1..EN in completion order to every completed investigation shell/MCP observation, including errors, empty responses and truncation, with no relevance filtering, deduplication, summarization or source supplementation. Serena initial_instructions is the single predeclared setup/manual exception: preserve it separately and charge its overhead. Draft statements never become new packet evidence. An observation record retains the exact raw CLI event bytes and their hash; the review sees its decoded command/arguments and aggregated_output or result/error without content edits. This is explicitly CLI-rendered observation evidence, not guaranteed original model function arguments or complete model input bytes. All cited source lines must be authentic and visible in the supplied observations.

Freeze the review serializer before experimental task dispatch. Use tiktoken0.12.0/o200k_base, with encoder-cache SHA256 recorded in TOKENIZER.json, for reference packet segmentation. Retain the60,000-token returned-packet limit under that frozen reference tokenizer, and disclose that it is not a provider input-count endpoint. No trimming to fit. Oversize packets, empty/refused/timed-out investigations and model budget failures retain a failed workflow with review-not-dispatched and R_sys=U=0, critical completion failure and failed source/protocol gates. A nonempty wrong draft proceeds unchanged to review. Invalid final JSON is a scored failure. Provider misses and honest unsupported results after a valid setup are outcomes, not infrastructure excuses.

Provider order uses each of the six permutations once across T1..T6 in repetition1, and reversed corresponding permutations in repetition2. Serial execution prevents overlap between measured workflows. Cold install/config/index/readiness/refresh, source/runtime hashing, prompt assembly, review assembly and validation all count. Report tool/model-session durations and complete workflow duration separately. Repository/provider immutable indexes may be reused after declared first setup; agent/server sessions remain fresh, no task-answer memory persists. Freeze first-task order and cache-state policy; do not claim OS filesystem cache purging. Charge cold setup to the first workflow and show it separately; no hidden amortization. Also retain suite totals and one-time research preparation separately.

Unchanged quality and latency criteria

Freeze one repository-answerable obligation denominator per task, shared by all three arms. R_sys is correctly resolved answerable obligations divided by that denominator; retrieval omissions cannot shrink it. Honest uncertainty receives U credit only for genuinely repository-unanswerable requirements. Missing requested behavior, actions or justified verification loses coverage. Semantic equivalents receive equal credit; a preferred qualified symbol spelling is not a requirement unless behavior depends on it.

Report all runs, task medians, macro task-median means, repository summaries and paired differences. For12 Madar workflows: zero critical wrong/unsafe claims, critical obligation failures, false source citations or unsupported concrete source assertions; every task-median R_sys>=0.80; macro task-median R_sys>=0.90; every required U case correct. Apply the same scorecards to comparators. At most0.05 macro-quality loss versus Serena; no critical quality loss versus Native. These preserve the absolute/system gates, not the old fixed-evidence mechanism-gain threshold.

For elapsed time and separately for the declared reference usage index: sum(Madar task medians)/sum(Native task medians)<=0.85; versus Serena<=1.05. No task-median ratio>1.15 and no paired individual ratio>1.25 against either comparator. Include complete investigation+review+attributable setup/assembly/validation. Report investigation-only and review-only numbers too. Failure denominators remain; a failed workflow cannot earn an efficiency pass merely by ending quickly. All12 Madar quality/protocol workflows must pass before an efficiency advantage is accepted.

Subscription accounting and finite limits

Actual attributable dollars and exact subscription-quota consumption are UNAVAILABLE, not zero or passed. The old$30 actual-charge gate and worst-case predispatch reservation are NOT established; they are not admission requirements for this explicitly different subscription protocol. No claim of monetary savings follows.

For a separate reference usage index, retain the historical#741 weights:4uncached_input +0.40cached_input +20*output, divided by1,000,000. Uncached=input-cached exactly once; reasoning is included through reported output, never added twice. This frozen index is not a current tariff, invoice or subscription charge. Preserve reported cache-write values without inferring invoice semantics; CLI may default zero. Missing input/cache/output semantics or missing terminal usage prevents an efficiency verdict. No new pricing/account lookup is required to use the fixed reference index.

Retain90 model sessions total including preparation, graders and any replacements;72 planned investigation/review sessions;8hours campaign wall time;3,000,000 cumulative input and600,000 output tokens. Keep per-experimental-session90,000 cumulative input and8,000 output as POST-RESPONSE admissibility checks. Each complete workflow<=12minutes and60 ordinary+MCP operations, investigation<=8minutes, review<=4minutes and0operations; cold setup consumes the12-minute allowance. Local cancellation is reactive and cannot guarantee cancellation of backend work. A completed session's reported overrun is retained as a failure; stop subsequent dispatch on campaign limit exhaustion. Never describe these as hard in-flight token/cost bounds.

Book a unique serial session ID before starting each request, keep starts/finishes/interruptions and cumulative measured usage in an append-only run register. No concurrent model requests or silent retries. Uncertain/missing usage stops further dispatch with its consumed session retained. Reserve remaining session slots for planned investigations/reviews before auxiliary work; preparation/adjudication must fit remaining capacity. At most one verified infrastructure-invalid complete workflow may be replaced under identical task/source/config, preserving and charging both attempts; stop on a second infrastructure invalidity. Model/protocol/provider failures receive no free replacement. No partial sample is a passed campaign.

Grading and disposition

Two fresh source-based graders independently rate anonymized final answers against frozen truth before seeing provider labels, costs or prior scores. Provider/tool names in grading metadata are suppressed; source/claim text remains unchanged, so imperfect stylistic blinding is disclosed. Use stable random anonymous IDs with a separately held map. Grade R_sys, U, critical completion, every concrete assertion/citation and proposed verification; a protocol checker separately evaluates format and observation integrity. Batch grading by repository to stay within session capacity. A third fresh adjudicator resolves disagreement only against the frozen oracle when capacity permits; no oracle rewriting. Unresolved adjudication or defective truth stops a campaign verdict.

Report three separate scorecards: acquisition/source evidence, final answer/plan quality, and complete time/reference usage. If Madar satisfies all applicable gates, the supported result is an advantage on this small sample under these configured agent systems and observed measurements. Actual-dollar advantage remains unestablished. If a gate fails, retain the diagnostic location from source -> returned evidence -> final use; do not lower thresholds, expand the sample, rerun for acceptance, reopen a code mechanism or create an architecture issue automatically. Recommend the provider/process supported by the results. #740 remains open until the product decision and resulting work disposition are explicit; no merge or release is implied.

Frozen one-task diagnostic successor protocol

Madar one-task diagnostic v2

Authorized by the maintainer's proceed after the stopped v1 report. This successor corrects one source-inaccurate expected-answer obligation through Codex CLI, then compares Native, Madar0.32.1 core and pinned Serena on the corrected existing Got T3 task. No new task search, runner implementation or product edit. The stopped v1 and its229-file manifest remain immutable. This is one task, one repetition and three workflows, not the six-task campaign and not evidence that its gates passed. T3 was chosen prospectively because it is the task being repaired; no scored provider answers exist. Coordinator knows the previous rejection, is not an independent task owner and will not semantically grade responses.

Correction: one fresh CLI author receives T3, its original oracle and the rejecting source review. Return a source-justified data patch confined to T3-O2 wording and directly dependent alternatives/wrong-claim/verification fields. Preserve task text, obligation count, answerability and criticality. Root applies only the returned patch mechanically to a new oracle version and verifies old-value equality, unchanged other fields and source hashes. A separate fresh source reviewer receives only T3, the revised oracle and source, with no prior verdict or desired winner; it admits or rejects the pair. Rejecting that pair ends this successor without provider scoring. Original artifacts are never overwritten.

The existing recorder remains f636c000ab06449d5b971358c5db826a0c7adbac, Node22.22.3, Codex CLI0.149.0 requesting gpt-5.6-sol/medium/default. Use existing ChatGPT login only. No API account, key, endpoint/count/billing call or fallback. Fresh ephemeral investigator and review sessions use the recorder's source/config controls. Immutable returned weights, exact dollars, hard in-flight token bounds and subscription quota attribution remain unavailable. The same-user source/role isolation is observational, not an adversarial filesystem guarantee.

Got source687eb7dcc100ea3e548ebba227173b886789e670 and installed provider/runtime/compiler versions remain unchanged. Infrastructure adviser may perform model-free initialization/tools-list checks and record exact setup instructions, advertised schemas, source hashes, timing and shutdown; no task/answer access or source-query tools. Coordinator/adviser chat usage is outside the CLI register and unmeasured, as in the earlier scope disclosure. Do not call it total account usage. The infrastructure role never curates, investigates or grades this task.

Freeze corrected oracle, exact task, prompts, runtime/config/schema hashes and order before investigators. Use original T3 repetition1 order: Madar, Native, Serena. One repetition cannot remove order/cache effects. Reuse existing static indexes; servers and agent sessions are fresh. Report previously measured cold install/index preparation separately, actual readiness work separately and run-stage investigation/review/assembly/validation elapsed separately. Present a constructed cold-cost scenario only if components and reuse are explicit; do not claim an observed three-system cold-speed ranking from this warm diagnostic.

Every investigator receives the same conditional investigation instruction from experiment741/PROMPTS.md, the task and explicit source/tool/budget boundaries. Provider guidance remains neutral: use when relevant, verify and fall back to ordinary source reads. Each investigator has8minutes, at most6 ordinary/MCP operations, and the same request to batch narrow reads and keep output concise. Native has no MCP; Madar/Serena have the existing exact seven server tools. Generic resource calls are observable protocol failures, never an allowed extra tool. This is observable rejection, not prevention-before-access.

Every nonempty completed investigation draft proceeds unchanged to the identical ordinary review A, with all returned observations serialized by the v1 frozen method and no new source acquisition. Review has4minutes and0operations. Packet reference limit60000 under tiktoken0.12.0/o200k_base; no trimming. Preserve every failed tool call, error, empty result and truncation. Serena initial_instructions remains the sole setup exception. Invalid/refused/timed-out/empty investigations or oversized packets are reported as failed workflows; never hide them or select another task. JSON/schema/link and actual observed-source support checks remain required.

Report90k cumulative input/8k output per-session overruns explicitly. In this exploratory successor they are resource-limit flags, not grounds to discard an otherwise nonempty draft or conceal its answer quality. This differs explicitly from v1's scored-failure/no-review policy; no v1 qualification is claimed. Total combined allowance is unchanged: previous2375165 input and67511 output, plus this successor, must remain below3million input/600000 output before further dispatch. Previous10 booked slots carry forward toward90; all new correction/check/investigation/review/grader requests book unique serial slots and preserve terminal counters. Retain8hours from the prior campaign start06:07:53 UTC. Limits are reactive post-response checks; missing/uncertain usage or exhaustion stops further dispatch, with partial evidence reported as incomplete. No silent retries or additional samples. Plan10 new sessions: correction, source confirmation,6 investigator/review sessions,2 graders. Remaining allowance is624835 input/532489 output; it is a ceiling, not a guaranteed reservation.

Two fresh blinded graders receive anonymized final answer/observation packets, the corrected frozen oracle and source; no provider labels, timing, costs or previous scores. Grade the shared repository-answerable denominator, genuine uncertainty, each critical obligation, concrete source claims/citations, and proposed verification. Provider names in metadata are suppressed while claim/source text remains unchanged; stylistic blinding is imperfect. Record each independent score and any disagreement; no extra adjudicator in this diagnostic. Root computes arithmetic only. Report acquisition, final quality and time/reported usage separately. If graders disagree materially, report the disagreement rather than choose a favorable score. No single-task result establishes an overall winner, release readiness, monetary savings or a previously failed historical gate.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    architectureCross-cutting design or substrate decisionspriority:p1roadmapTracked on the public README roadmap

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions