Repository navigation
fix(core): keep every tool result paired when a tool-call id collides [K-03] - #6
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Close the two ways a tool-call id collision loses work in one round:
<id>_d<n>) before dispatch, so every tool result keeps its own pairing id.400.Why
A. Collisions drop results.
core.py:1104(pre-change) coalesced the streaming id withcall_id = stc["id"] or stc["name"]. Nothing made the ids distinct, but the id is the key of every downstream result map:completed_results[call["id"]] = (res, t_elapsed)(core.py:1268),blocked_results[call["id"]](core.py:1204), and thetool_call_idof each written tool message (core.py:1289). Two calls in one batch sharing an id therefore overwrite each other — one tool's result is lost, the other's is replayed for both — and the model gets tworole:"tool"rows with the sametool_call_id, which strict providers answer with a 400.Measured on the parallel dispatch path with two
read_filecalls sharing idcall_dup(scratch probe, real_run_conversation_loop):a.py's result never reached the model at all.B. The 400 recovery fires on the wrong error.
core.py:975(pre-change) matched"tool_calls" in msg or "400" in msg, set the one-shot_stream_retried, ran the destructive_drop_dangling_tool_calls()and retried — regardless of whether anything was dangling, and before the key-rotation/backoff branches below it. An unrelated error whose text contains400(a token count like "maximum context length is 4000 tokens", or "… 400 requests/min") was therefore stolen from the rate-limit recovery path and retried with no backoff, while a legitimate round was removed from history. Same scenario, real loop, onmain: 5 stream attempts before the turn gave up; with the fix, 1 and the error is surfaced.What changed
utils/tool_ids.py(new):uniquify_tool_call_ids(calls) -> int. Renames later duplicates to<id>_d<n>(n from 2, first occurrence wins), mutating in place, returning the number renamed. Deterministic on purpose — these ids ride the prompt-cache prefix, so a random suffix is not acceptable. Blank / non-string ids are skipped.core.py: the round assembly buildscalls, uniquifies their ids, then keysinvalid_argsoff the final ids (so a vetoed call and thetool_call_idwritten for it can no longer disagree). Logs a warning when a rename happens.core.py: the 400 branch calls_drop_dangling_tool_calls()first and only consumes_stream_retried+ retries when it returned > 0 messages; otherwise it logs and falls through to the existing rate-limit / overloaded recovery.Provenance
hermes:agent/message_sanitization.py:531 uniquify_tool_call_ids— same_d<n>scheme and the same determinism constraint ("never uuid4 — these ids feed prompt-cache prefixes");hermes:run_agent.py:1220-1241for the id policy. Re-implemented for koza's dict-shaped calls with koza names/logging; no upstream code or identifier was copied.Verification
tests/test_tool_call_ids.py(new, force-added —tests/is gitignored): 9 tests, 3 of them driving the real_run_conversation_loopwith a stubbed streaming provider. Proven to catch the bugs — withcore.pyrolled back tomainand the new tests in place:and the restored tree is
9 passed. Scratch probe (~/.hermes/cache/scratch/k03_probe.py, not committed) produced the before/after output quoted above and confirmed the genuine dangling-400 path still repairs and retries (rounds_started == 2).Risk / rollback
Low. No new dependency, no config key, no signature change outside
core.py's loop. The id rename only triggers when the provider actually reused an id, so the common case is byte-identical on the wire (no prompt-cache change). Reverting the singlecore.pyhunk plus droppingutils/tool_ids.pyrestores prior behaviour.Roadmap: K-03