fix(core): re-prompt when finish_reason=tool_calls arrives with no call buffered [K-04] - #7
Merged
Merged
Conversation
…ll buffered [K-04]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
When a provider ends a round with
finish_reason="tool_calls"but streams no tool-call chunk at all, the turn's narration text is no longer accepted as the final answer. Koza now re-prompts for the actual call, bounded to 3 consecutive stalls.Why
On
main,core.py:1110-1113collapsed "no buffered call" and "the model finished talking" into the same path:_tool_bufis empty whenever the provider never emitted a__tool_chunk__— which includes the case where_finish_reasonwas set to"tool_calls"(core.py:935-936, set from the API'schoice.finish_reasoninproviders/openai_provider.py:92,providers/deepseek_provider.py:208). This is the classic narration-only stall: the model writes "I'll read a.py now…", the provider (or an interrupt mid-retry) closes the stream withtool_callsand an empty array, and koza appends that plan text as the assistant's final message and returns. The task stops half-done, silently, with a plan presented as the answer.Reproduced against
main(core.pyunpatched), one narration round then the real call then the answer:One round, no tool ever ran, plan delivered as the answer.
What changed
core.pyonly (+47/-1):_DROPPED_TOOLCALL_NUDGE(core.py:427) — the re-prompt text ("Your previous turn indicated a tool call but none was included. Do not narrate a plan or restate intent — issue the actual tool call now…").if not _tool_buf:branch (core.py:1119-1153) now first checks_finish_reason == "tool_calls". If so and fewer than 3 consecutive stalls have happened, it appends the narration + the nudge (both tagged_dropped_toolcall_nudge) andcontinues instead of returning. After the budget is spent the narration is delivered as before, so a model that only ever narrates still terminates._dropped_toolcall_retriesis a loop-local (core.py:900), reset whenever a batch actually reaches execution (core.py:1251) and on a genuine text turn (core.py:1151), so the bound counts consecutive stalls.finallycleanup (core.py:1487) now strips_dropped_toolcall_nudgemessages as well as_empty_recovery_synthetic, so the retry scaffolding never lands in the durable transcript.Provenance
Mechanism ported from the upstream Hermes agent:
agent/conversation_loop.py:975-979(_DROPPED_TOOLCALL_NUDGE_CONTENT) and its consumeragent/turn_final_response.py:279-311(bounded re-prompt, marked ephemeral scaffolding). Re-implemented in koza's own style — koza name, koza cleanup convention, no imported identifiers.Verification
Real output from this session (
.venvactive):tests/test_dropped_tool_calls.py(new, force-added —tests/is in.gitignore): 6 tests, 3 of them driving the real_run_conversation_loop. Test power,core.pyrolled back tomainviagit stash push -- core.py:Scratch probe:
~/.hermes/cache/scratch/k04_probe.py(not committed). The bounded-termination case confirms 1 round + at most 3 re-prompts, after which the narration is delivered — no infinite loop (the outerMAX_ROUNDS = 25also caps it).Risk / rollback
Low. The change only fires when
finish_reason == "tool_calls"and nothing was buffered — a statemaincurrently ends the turn on; every other path is unchanged (a text turn withfinish_reason="stop"orNonestill returns on the first round, covered by two regression tests). Rollback: revert the commit; behaviour returns tomain.Roadmap: K-04