Repository navigation
fix(dsh): recover Bonsai sessions and correct decode timing - #1
Merged
Merged
Conversation
Repeated unchanged tool calls could run indefinitely, while buffered tool arguments made the client report impossible generation rates. Publish Engine phase timings in Responses and rebuild the local timing projection from valid measurements. Bound duplicate execution, preserve its evidence across resume and compaction, and allow one recovery attempt with obsolete reasoning removed from the projected prompt. Reserve answer space with exact prompt counts, reclaim older reasoning when necessary, and recover nonconvergent compaction with bounded whole source excerpts backed by the complete indexed archive. Keep durable history and user/tool evidence intact. Select low thinking, temperature 0.2, RK4 KV and FP32 recurrent state for the 64K profile, using draft-head memory for cache precision rather than reducing context capacity. Include the deployable preset, independent coding oracles and live evaluators. The three coding fixtures passed 315 cases after repair, and the long-context fixture completed at 58,543 prompt tokens. Detached replays of the failing session recovered both its next operation and compaction without executing its proposed file-analysis command. These results qualify the exercised workflows, not arbitrary first-draft accuracy. Validation: 67 offline DSH regressions passed on Node 24.19.0 and DSH 0.1.5-rc.3. The host-only Responses suite passed with MSVC 14.44 and the existing CUDA 13.1 libraries. Live qualification used the RTX 3080 10 GB; the deployed preset matched the repository and the server retained its 65,536-token capacity. git diff --check passed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Repeated unchanged tool calls could run indefinitely, while buffered tool arguments made the client report impossible generation rates. Publish Engine phase timings in Responses and rebuild the local timing projection from valid measurements. Bound duplicate execution, preserve its evidence across resume and compaction, and allow one recovery attempt with obsolete reasoning removed from the projected prompt.
Reserve answer space with exact prompt counts, reclaim older reasoning when necessary, and recover nonconvergent compaction with bounded whole source excerpts backed by the complete indexed archive. Keep durable history and user/tool evidence intact. Select low thinking, temperature 0.2, RK4 KV and FP32 recurrent state for the 64K profile, using draft-head memory for cache precision rather than reducing context capacity.
Include the deployable preset, independent coding oracles and live evaluators. The three coding fixtures passed 315 cases after repair, and the long-context fixture completed at 58,543 prompt tokens. Detached replays of the failing session recovered both its next operation and compaction without executing its proposed file-analysis command. These results qualify the exercised workflows, not arbitrary first-draft accuracy.
Validation: 67 offline DSH regressions passed on Node 24.19.0 and DSH 0.1.5-rc.3. The host-only Responses suite passed with MSVC 14.44 and the existing CUDA 13.1 libraries. Live qualification used the RTX 3080 10 GB; the deployed preset matched the repository and the server retained its 65,536-token capacity. git diff --check passed.
Problem and scope
Related Issue:
Implementation
Verification