Skip to content

fix(dsh): recover Bonsai sessions and correct decode timing - #1

Merged
ahnafnafee merged 1 commit into
masterfrom
fix/bonsai-session-recovery
Sep 27, 2026
Merged

ahnafnafee merged 1 commit into
masterfrom
fix/bonsai-session-recovery

Conversation

@ahnafnafee

Copy link
Copy Markdown
Owner

Repeated unchanged tool calls could run indefinitely, while buffered tool arguments made the client report impossible generation rates. Publish Engine phase timings in Responses and rebuild the local timing projection from valid measurements. Bound duplicate execution, preserve its evidence across resume and compaction, and allow one recovery attempt with obsolete reasoning removed from the projected prompt.

Reserve answer space with exact prompt counts, reclaim older reasoning when necessary, and recover nonconvergent compaction with bounded whole source excerpts backed by the complete indexed archive. Keep durable history and user/tool evidence intact. Select low thinking, temperature 0.2, RK4 KV and FP32 recurrent state for the 64K profile, using draft-head memory for cache precision rather than reducing context capacity.

Include the deployable preset, independent coding oracles and live evaluators. The three coding fixtures passed 315 cases after repair, and the long-context fixture completed at 58,543 prompt tokens. Detached replays of the failing session recovered both its next operation and compaction without executing its proposed file-analysis command. These results qualify the exercised workflows, not arbitrary first-draft accuracy.

Validation: 67 offline DSH regressions passed on Node 24.19.0 and DSH 0.1.5-rc.3. The host-only Responses suite passed with MSVC 14.44 and the existing CUDA 13.1 libraries. Live qualification used the RTX 3080 10 GB; the deployed preset matched the repository and the server retained its 65,536-token capacity. git diff --check passed.

Problem and scope

Related Issue:

Implementation

Verification

Repeated unchanged tool calls could run indefinitely, while buffered tool
arguments made the client report impossible generation rates. Publish Engine
phase timings in Responses and rebuild the local timing projection from valid
measurements. Bound duplicate execution, preserve its evidence across resume
and compaction, and allow one recovery attempt with obsolete reasoning removed
from the projected prompt.

Reserve answer space with exact prompt counts, reclaim older reasoning when
necessary, and recover nonconvergent compaction with bounded whole source
excerpts backed by the complete indexed archive. Keep durable history and
user/tool evidence intact. Select low thinking, temperature 0.2, RK4 KV and
FP32 recurrent state for the 64K profile, using draft-head memory for cache
precision rather than reducing context capacity.

Include the deployable preset, independent coding oracles and live evaluators.
The three coding fixtures passed 315 cases after repair, and the long-context
fixture completed at 58,543 prompt tokens. Detached replays of the failing
session recovered both its next operation and compaction without executing
its proposed file-analysis command. These results qualify the exercised
workflows, not arbitrary first-draft accuracy.

Validation: 67 offline DSH regressions passed on Node 24.19.0 and DSH
0.1.5-rc.3. The host-only Responses suite passed with MSVC 14.44 and the
existing CUDA 13.1 libraries. Live qualification used the RTX 3080 10 GB;
the deployed preset matched the repository and the server retained its
65,536-token capacity. git diff --check passed.
@ahnafnafee
ahnafnafee merged commit 8d95954 into master Sep 27, 2026
3 of 4 checks passed
@ahnafnafee
ahnafnafee deleted the fix/bonsai-session-recovery branch September 27, 2026 13:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant