Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
55 commits
Select commit Hold shift + click to select a range
3f639df
chore(release): open dev at 2.58.0 before releasing 2.57.0 (#4827)
github-actions[bot] Sep 16, 2026
406606a
docs(devlog): record the 2.57.0 release and its registry propagation …
lidge-jun Sep 16, 2026
e2304ce
fix(lib): wait for a killed subprocess to die before abandoning it (#…
lidge-jun Sep 16, 2026
2941f95
fix(responses): enforce the shutdown drain reserve on the injected sp…
lidge-jun Sep 16, 2026
ade8155
test(budget): restore the 45s spawn signal and stop hardening a dispo…
lidge-jun Sep 16, 2026
6ed8986
test: fix the causes behind two disabled test families and re-enable …
lidge-jun Sep 16, 2026
5d98281
docs(devlog): correct the 2.57.0 Windows evidence and record the afte…
lidge-jun Sep 16, 2026
8c59174
ci: fail on the first timeout or Bun crash instead of retrying to gre…
lidge-jun Sep 16, 2026
cfdf5c3
perf(windows): stop hardening every secret temp twice (#4836)
lidge-jun Sep 16, 2026
abae31a
fix(lib): let a killed child actually die before anyone deletes what …
lidge-jun Sep 17, 2026
7ef3f67
test: assert these contracts by behaviour instead of by source shape …
lidge-jun Sep 17, 2026
19bdcaa
ci(windows): restore the margin the six-shard leg lost, and make a br…
lidge-jun Sep 17, 2026
121405b
test(codex): isolate the lock child's database and wait on a real sig…
lidge-jun Sep 17, 2026
7a29e7b
test(responses): assert the stream completed, not that it lacks the d…
lidge-jun Sep 17, 2026
e18ca24
perf(cli): stop config show from importing the connect graph to read …
lidge-jun Sep 17, 2026
7868f5d
test(web-search): drive the connect deadline on a virtual clock (#4867)
lidge-jun Sep 17, 2026
f1dfda8
ci(windows): the batch leg is one test run, so it must not queue agai…
lidge-jun Sep 17, 2026
b25e449
docs(devlog): audit the native control stack as one integration contr…
lidge-jun Sep 17, 2026
58b8400
docs(devlog): plan the L6 contract and quality fix lane (#4879)
lidge-jun Sep 17, 2026
a1fe84b
docs(devlog): open the L4 Responses private-field and history-repair …
lidge-jun Sep 17, 2026
09f0421
docs(providers): say where transientRetryOn5xx does not apply (#4894)
lidge-jun Sep 17, 2026
5ea488f
fix(adapters): charge adapter-owned retry sends to the request send b…
luvs01 Sep 17, 2026
484dcf5
docs(devlog): plan the L3 retry, admission and combo-recovery unit (#…
lidge-jun Sep 17, 2026
27e4310
fix(bridge): classify clean terminal message phases (#4885)
lidge-jun Sep 17, 2026
4b234da
fix(responses): stop forwarding Codex's private access_programs to th…
lidge-jun Sep 17, 2026
e05ae26
ci: gate the docs site build on pull requests (#4895)
lidge-jun Sep 17, 2026
863d68d
docs(devlog): record the L4 destination-boundary reversal and both re…
lidge-jun Sep 17, 2026
1ac246c
fix(providers): support Z.AI model discovery contract (#4886)
lidge-jun Sep 17, 2026
6b6d6e6
fix(claude): report confirmed first-frame usage (#4891)
lidge-jun Sep 17, 2026
dfa02fe
fix(codex): restore routing without waiting on paginated history (#4896)
lidge-jun Sep 17, 2026
ccf3b03
fix(codebuddy): refuse bare DSML tool names (#4887)
lidge-jun Sep 17, 2026
f4b52ee
fix(cursor): quarantine textual TOOL_CALL markers off the text channe…
lidge-jun Sep 17, 2026
25311bc
fix(providers): refresh the Alibaba Token Plan catalogs against gatew…
lidge-jun Sep 17, 2026
cdd564b
fix(models): mirror model capacity onto the /v1/models top level (#4889)
lidge-jun Sep 17, 2026
fb0926e
fix(cursor): prefer observed checkpoint maxTokens as the overflow cei…
lidge-jun Sep 17, 2026
b60a628
fix(codex): close pool eligibility inside the caller-owned preview re…
lidge-jun Sep 17, 2026
6d19a07
fix(deps): bump hono, astro, and docs-site overrides to resolve audit…
agentHits Sep 17, 2026
4971cdf
fix(settings): apply the Desktop switches and report effective state …
lidge-jun Sep 17, 2026
a50f1de
test(windows): report a dead child instead of an opaque readiness tim…
lidge-jun Sep 17, 2026
3bdaf61
fix(models): reject an invalid custom-model context window instead of…
lidge-jun Sep 17, 2026
ee28833
fix(cursor): remint conversation after incomplete tool-call streams (…
MerryEcho Sep 17, 2026
7c8e961
fix(xai): seed Responses tool-result adjacency for interrupted Codex …
MerryEcho Sep 17, 2026
eca65bd
test(budget): size the competing-OFF bounds on the preparation window…
lidge-jun Sep 17, 2026
46b3816
fix(ollama-native): defer a mid-turn conversational message instead o…
briascoi Sep 17, 2026
519db59
feat(responses): preserve native mid-turn steering over WebSocket (#4…
luvs01 Sep 17, 2026
ef56be8
fix(combos): allow single-target combo with waitForCooldownMs to retr…
agentHits Sep 17, 2026
41f1832
fix(combo): fail over zero-output SSE errors (#4817)
87003697 Sep 17, 2026
9052ddf
test(combos): hold the zero-output bare-error case in a sibling file …
lidge-jun Sep 17, 2026
3ec1af6
feat(responses): relay native multi-agent function-result injection (…
luvs01 Sep 17, 2026
4fce2c7
docs(devlog): open the 2.58.0 release train (#4909)
lidge-jun Sep 17, 2026
fa26404
feat(responses): extend native result continuations and preserve host…
luvs01 Sep 17, 2026
f671934
fix(responses): bound steering confirmation waits and preserve sparse…
lidge-jun Sep 17, 2026
5061f2c
feat(responses): support safe steering settings, public API transport…
lidge-jun Sep 17, 2026
6fe4cd0
Merge pull request #4915 from lidge-jun/release/2.58.0
lidge-jun Sep 17, 2026
3a20c3d
Merge remote-tracking branch 'origin/main' into release/preview-2.58.0
lidge-jun Sep 17, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
396 changes: 271 additions & 125 deletions .github/workflows/ci.yml

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

33 changes: 33 additions & 0 deletions devlog/_plan/260916_cursor_http2_toolcall/000_plan.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# 260916 — Cursor HTTP/2 toolcall complementary stack

Textual `[TOOL_CALL]` markers still reach Codex as assistant text after #2305
renamed the display alias. That leak is the first stacked PR. Observed
`tokenDetails.maxTokens` is the second. No local suite; hosted CI verifies.

## Loop spec

- Loop archetype: satisfy-spec complementary stack.
- Trigger: user asked to fill Cursor HTTP/2 gaps (Aside + senpi/shunt/jcode)
with toolcall leak as the primary cause, then stack-PR push `--no-verify`
without running the local suite.
- Goal: 2 stacked PRs on `dev` (manual chain, not GitHub native stacks).
- Non-goals: `cursor-agent` subprocess, cursor-proxy wholesale, native-exec
default-on, host-credential import, `bun test` / `bun run test`, merge.
- Verifier: GitHub PR URLs + `gh pr view --json baseRefName,headRefName`.
Local suite: NOT RUN (user restriction).
- Stop: both PRs exist with parent/child bases.
- Memory artifact: this directory.
- Terminal: DONE (PRs opened) / BLOCKED (push/template) / UNSAFE (exec default-on).
- Shipped (2026-09-16, `gh pr view` bases):
- L1 https://github.com/lidge-jun/opencodex/pull/4815 `dev` ← `cursor/l1-text-toolcall-quarantine`
- L2 https://github.com/lidge-jun/opencodex/pull/4816 `cursor/l1-text-toolcall-quarantine` ← `cursor/l2-observed-max-tokens`
- Escalation: live `api2` vs `agentn.global.api5` host cutover.

## Work-phase map

1. wp0 — this roadmap (docs-only).
2. wp1 / 010 — L1 quarantine + promote textual tool calls.
3. wp2 / 020 — L2 persist observed `maxTokens` for the overflow size prior.

Host URL migration (`pleaseai/shunt` → `agentn.global.api5.cursor.sh`) stays
OUT until a live 464/ALPN failure is recorded against current OpenCodex pins.
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# 010 — L1 textual toolcall quarantine

## IN

- NEW `src/adapters/cursor/text-toolcall.ts`
- MODIFY `src/adapters/cursor/protobuf-events.ts` (`textDelta`, finalize, state)
- MODIFY `tests/providers/cursor/cursor-protobuf-events.test.ts` (#2305 block)
- MODIFY `structure/providers/cursor.md`

## OUT

Host cutover, stall-resume, CLI spawn, new test-layout file.

## Diff contract

- `drainCursorTextToolCalls(pending, chunk)` extracts complete
`[TOOL_CALL]name[ARGS]{json}` blocks, folds `mcp_opencodex-responses_*`
names, returns surrounding prose + pending opener.
- Incomplete markers hold up to 64 KiB then drop (no leak).
- Advertised names → atomic `tool_call_start/delta/end` via existing
`recordToolCall` + `commitToolCall`.
- Unadvertised names: strip only.
- Finalize deletes `pendingTextToolCall`.

## Accept

- Marker + surrounding prose: text has no `[TOOL_CALL]`, tool events exist.
- Split deltas promote on the second chunk.
- Finalize of a held opener emits `done` without the marker.
- Activation: `textDelta` containing a complete or split marker.
Observable: no `[TOOL_CALL]` in mapped events; advertised name becomes a
committed tool call.

## Verifier

NOT RUN locally. Hosted `bun test tests/providers/cursor/cursor-protobuf-events.test.ts`.

## Shipped

https://github.com/lidge-jun/opencodex/pull/4815
`dev` ← `cursor/l1-text-toolcall-quarantine` (`407bf3ce56`).
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# 020 — L2 observed maxTokens ceiling

## IN

- MODIFY `src/adapters/cursor/protobuf-events.ts` checkpoint handler
- MODIFY `src/adapters/cursor/discovery.ts` `inferCursorContextWindow`
- MODIFY `src/adapters/cursor.ts` `cursorRequestSizeContext`
- MODIFY `src/adapters/cursor/cursor-errors.ts` callers if the size prior
needs the observed window
- Tests in existing `tests/providers/cursor/cursor-errors.test.ts` /
`cursor-protobuf-events.test.ts`

## OUT

Disk store (`cursor-context-limits.json`), senpi admission amputation,
client-version bump.

## Diff contract

- On `conversationCheckpointUpdate`, record positive `tokenDetails.maxTokens`
per wire model id in a process-local map (same shape as usage carry-forward).
- `inferCursorContextWindow(modelId, observed?)` prefers a positive observed
ceiling, else today's id heuristic.
- `cursorRequestSizeContext` feeds that window into the existing 0.5-window
overflow vs 429 prior.

## Shipped

Process-local map in `discovery.ts`; checkpoint records a positive
`maxTokens` when `wireModelId` is set from `live-transport.ts`.
`inferCursorContextWindow(modelId, observed?)` prefers explicit then
recorded then heuristic. `cursorRequestSizeContext` is unchanged except
the comment — it already calls `inferCursorContextWindow`.

https://github.com/lidge-jun/opencodex/pull/4816
`cursor/l1-text-toolcall-quarantine` ← `cursor/l2-observed-max-tokens` (`8763dee2d2`).

## Accept

- Checkpoint with `maxTokens: 32000` makes a 20-token request classify as 429
(small vs observed window), not overflow.
- Zero/missing `maxTokens` keeps the heuristic (first checkpoint is 0 on senpi).
- Activation: checkpoint frame with `maxTokens > 0`, then a bare
`resource_exhausted` on a small payload.
Observable: `classifyCursorError` stays on the 429 class.
84 changes: 84 additions & 0 deletions devlog/_plan/260917_2570_release_train/040_release.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,84 @@
# wp5 — the 2.57.0 release

Closed. 2.57.0 is on `main` and `preview`, the GitHub release exists, and `npm publish` returned
success with a signed provenance statement. Registry propagation is tracked at the end.

## Sequence, with evidence

| Step | What happened | Evidence |
| --- | --- | --- |
| Freeze the candidate | `1831193294` on `dev`, `package.json` 2.57.0 | Cross-platform CI push run `35131181996`: success. First green `dev` run since `35091966777`. **Windows was skipped, not run** — see the correction below. |
| Move `dev`'s version line first | #4827 opened by `dev-version-bump.yml` run `35127386565`, merged as `3f639dfdad` | CI run `35127440220` success after rerunning a cancelled `macos 1/2`; Service lifecycle `35127440065` success. |
| Promote to `main` | #4829 merged as `44de45dfdc` | Merge commit, matching the 2.56.0 promotion #4694. `enforce-target` red by design. |
| Prove the release SHA | `44de45dfdc` | Cross-platform CI run `35133242171`: success. Service lifecycle run `35133242154`: success. Windows was skipped here too; the shards were run afterwards, below. |
| Dispatch `release.yml` | version 2.57.0, tag latest, dry-run false, `expected-sha=44de45dfdc33d30af22502d2bed98014fe16d83b` | Run `35135131119`: success. `Publish` step ends `+ @bitkyc08/opencodex@2.57.0`; provenance in the sigstore transparency log at logIndex 2865732791. GitHub release `v2.57.0` created 18:35:07Z. |
| Promote to `preview` | #4831 merged as `b70f3d7fcb` | `git diff origin/main HEAD` empty; only `package.json`'s version line conflicted and was resolved to `main`'s 2.57.0, the same resolution #4698 used. |

## Two decisions worth recording

**The red `dev` was not a reason to stop.** Five consecutive failing runs looked like a regression
and were not; `010_dev_green.md` has the per-run forensics. The largest class was already fixed by
the candidate's own parent (#4821 pinning Bun back to 1.4.0), and the remaining class was a
45-second spawn budget that a Windows runner beat by 5.7 seconds (#4830). Aside research confirmed
Bun 1.4.2 is still the latest stable and no released version fixes that Windows crash class, so the
1.4.0 pin stays.

**CodeQL's "10 new alerts including 8 high severity" was reviewed rather than waived.** Eight carry
alert numbers already open on `main` and one is the same flow as main's #175 at a shifted line.
Exactly one is new — #183, `js/insufficient-password-hash` at `src/codex/account-label.ts:31` —
and it is a false positive, because that SHA-256 produces a log label for API-key selection, not a
password hash. The reasoning is on #4829.

## Registry propagation

`npm publish` succeeded at 18:34:37Z and npm answered "Your package is being processed and may take
a few minutes to become available." The workflow's own `Post-publish registry smoke` step then read
the registry six times without confirming, recorded `verification=pending`, and said in its summary:
*inspect the registry before announcing availability; do not republish this version.*

It took about eight minutes. `https://registry.npmjs.org/@bitkyc08%2fopencodex/2.57.0` answered 404
through 18:42 and then 200; the packument's `modified` moved to 2026-09-16T18:42:50.857Z and
`dist-tags.latest` reads 2.57.0. `npm view @bitkyc08/opencodex version` agrees.

The step is doing its job and its bounded read window is simply shorter than npm's worst-case
processing time. Nothing needs changing: the warning is accurate, it does not fail the release, and
it tells the reader exactly what to do instead of republishing. A pending verification here means
wait and re-read, not cut another version.

## Correction: Windows was never run on the candidate or the release SHA

The "all six Windows shards included" claim above was wrong, and it is worth saying plainly because
it is the sentence a future release would have trusted.

`platform-windows` is dispatch-only by design — `.github/workflows/ci.yml` gates it on
`github.event_name == 'workflow_dispatch'` — so on a `push` or `pull_request` event the six shards
are always `skipped`, and the aggregate `ci` check accepts a `skipped` producer as a pass. Both runs
cited above are push runs:

| Run | Head | `windows N/6` |
| --- | --- | --- |
| `35131181996` | `1831193294` (candidate) | skipped |
| `35133242171` | `44de45dfdc` (release SHA) | skipped |

The Windows evidence that existed at publication time was run `35134620067` on `1504caaa83` — the
#4825 lane head, not the candidate and not the release SHA. All six shards were green there, which is
why the claim felt true; it was attached to the wrong commit.

**The gap is now closed after the fact.** Dispatch `35139132889` ran `ci.yml --ref main -f lane=all`
at `44de45dfdc`, the exact commit `@bitkyc08/opencodex@2.57.0` was published from, and all six shards
passed individually:

```
windows 1/6=success windows 2/6=success windows 3/6=success
windows 4/6=success windows 5/6=success windows 6/6=success
```

So 2.57.0 is sound on Windows. What failed was the evidence discipline, not the release.

Two things follow, and both are being handled in the 2.58.0 cycle rather than here:

- The aggregate `ci` gate cannot tell "deliberately not requested for this event" from "was requested
and did not start", because it accepts every `skipped` result unconditionally. Making that gate
event-aware is the subject of the CI-integrity lane.
- A release must not be publishable without a Windows `lane=all` at the exact promotion SHA. The
publish checklist treated that dispatch as a step; nothing enforced it.
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# L1 — preview read fence and dependency audit

Two independent safety units that deliberately do not stack on any feature branch.
Each ends as its own pull request against `dev` with exact-head CI evidence.

| Unit | Subject | Artifact |
|---|---|---|
| A | Issue #4850 — caller-owned `thread_spawn` preview still reads physical main `auth.json` | `010_issue_4850_preview_pool_eligibility_fence.md` |
| B | PR #4873 — dependency audit overrides for `hono` and the docs-site toolchain | `020_pr_4873_dependency_audit_review.md` |

## Why they are separate

Unit A changes runtime credential-boundary behaviour in `src/codex/` and
`src/server/responses/`. Unit B changes only `package.json` and lockfiles and is
authored by an outside contributor. Putting them on one branch would make the
contributor's commit un-landable on its own and would drag a credential-boundary
review into a dependency bump.

## Verification posture

No local suite, typecheck, build, or install runs in this lane. Correctness is
argued statically from the source and the call graph, and confirmed by hosted CI
at the exact head of each pull request. That constraint is why unit A's completion
criteria are written as observable read counts rather than as "the right token was
eventually sent": a behavioural assertion that hosted CI can run is the only proof
available here, and it is the stronger one anyway.

## Boundaries

This lane does not merge, does not push to `dev`, and does not rebase without an
instruction. It does not widen timeouts, add retries, skip platforms, or mask a
failure to make CI green. Windows jobs are dispatch-only, so any change with
Windows impact is reported rather than dispatched here.

Unreleased security analysis belongs in `.tmp/`, never in this directory.
Both units here concern already-public material: #4850 is a filed public issue
with the call path in its body, and #4873's advisories are published GHSA records.
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
# Unit A — issue #4850: pool eligibility is outside the preview read fence

## The gap, stated precisely

`src/server/responses/request-prepare.ts` already computes the ownership fence.
`previewRequestScopedMainCredential` is the route ownership predicate ANDed with
`hasCallerCodexBearer`, exactly as final authentication validates it, and
`nativeMainReadsForbidden` ORs it with retained recovery and a draining selector.
Quota priming, entitlement discovery, and the denied-model cache all honour it.

`previewSelectionOptions` does not. It carries `nativeMainSelectionOnly` and the
uploaded-file retention bit and stops there. When that object reaches
`previewCodexAccountForRequest -> pickPriorityPreemption -> getEligiblePoolAccounts`,
`codexAccountUnusableReason` in `src/codex/account-usability.ts` finds no
`isMainAccountTokenLive` override and falls through to its default,
`isMainAccountCredentialUsable()`, which opens and parses the physical `auth.json`.

It happens twice because the same options object is used twice: once for the direct
preview in `prepareResponsesRequest`, and once inside the callback
`applySubagentModelFallback` invokes per candidate model. The post-decryption
recovery re-preview builds `recoverySelectionOptions` the same way and has the same
omission.

## What is and is not at stake

Not a token leak. Final authentication never selects the physical main credential
for a caller-owned request: it passes `isMainAccountTokenLive: () => preserveRequestOwnedMainPin`
into its own selection options, so main is either served as the caller's own
credential or scored `main_credential_unavailable` and dropped. ADR-0086 already
rejected reading the physical main token for identity.

What is at stake is that operator-main liveness, cached quota, and plan state can
enter the score that decides whether a subagent's model is rewritten, for a request
that owns its credential. A preview that scores main differently from the resolution
it exists to predict is a correctness defect on top of the boundary defect.

## Chosen direction

Use the existing `CodexAccountUsabilityOptions.isMainAccountTokenLive` seam, and
give it the same value final authentication gives it rather than a preview-only
constant.

That answers the open question in the issue review directly. `preserveRequestOwnedMainPin`
is not an arbitrary choice: it is the only value that makes the preview agree with
the resolution in both branches. When the operator has an effective manual main pin
with quota headroom, final authentication returns the caller-owned main context, so
the request really is served by main and the preview should score main eligible.
When there is no such pin, final authentication drops main from pool eligibility,
and the preview must drop it too. A hardcoded `true` would be wrong in the second
case, and a hardcoded `false` would be wrong in the first.

Every input to that predicate is config, policy, or in-memory runtime state —
`activeCodexAccountPinned`, `isEffectiveCodexAccountPinned`, `pausedCodexAccountIds`,
the in-memory quota score, and `matchesMainQuotaCredential`, which compares HMACs
against an observed-credential record held in `main-account-cache.ts`. Nothing in
it opens a file, which is what makes it usable on the fenced side.

To keep preview and final authentication from drifting apart again, the predicate
moves into one exported function in `src/codex/auth-context.ts` that both callers
use. Two copies of a fence is how this gap appeared in the first place.

## Edit set

| File | Change |
|---|---|
| `src/codex/auth-context.ts` | Extract `requestOwnedMainPinState` and call it from `resolveCodexAuthContext` |
| `src/server/responses/request-prepare.ts` | Pass the synthetic `isMainAccountTokenLive` in `previewSelectionOptions` and `recoverySelectionOptions`, scoped to the ownership flag |
| `tests/responses/responses-preview-main-read-fence.test.ts` | Assertions (a) and (b) below |

No new test file, so `layout.json` and `test-layout-expected.json` are untouched.
No file here is on the size-ratchet baseline. No `src/` area is created or removed
and no invariant test disappears, so `structure:check` has nothing to consume.

## Completion criteria

Deliberately stricter than "the right token was eventually sent", because that was
already true before the fix and the defect survived it anyway.

(a) A caller-owned `thread_spawn` performs **zero** `auth.json` reads across the
whole request, asserted on the unfiltered read counter rather than through the
denial-cache stack filter that currently hides these two reads.

(b) Ordinary main selection is unchanged. A request with no caller bearer still
reads the physical credential and still selects main when it is healthy, so the
fix cannot be satisfied by making main globally ineligible.

(c) The #3166 healthy main-pin behaviour survives: a caller-owned request under an
effective main pin is still previewed as main.

## Risk

The behaviour change is confined to requests where `previewRequestScopedMainCredential`
is true. For every other request the option is absent and `account-usability.ts`
takes the identical default branch it takes today.
Loading
Loading