Skip to content

fix(core): re-prompt when finish_reason=tool_calls arrives with no call buffered [K-04] - #7

Merged
haydarkadioglu merged 1 commit into
mainfrom
feat/K-04-empty-toolcall-buffer
Oct 4, 2026
Merged

haydarkadioglu merged 1 commit into
mainfrom
feat/K-04-empty-toolcall-buffer

Conversation

@haydarkadioglu

Copy link
Copy Markdown
Owner

Summary

When a provider ends a round with finish_reason="tool_calls" but streams no tool-call chunk at all, the turn's narration text is no longer accepted as the final answer. Koza now re-prompts for the actual call, bounded to 3 consecutive stalls.

Why

On main, core.py:1110-1113 collapsed "no buffered call" and "the model finished talking" into the same path:

# ── No tool calls → pure text response, done ─────────────────────
if not _tool_buf:
    self.messages.append({"role": "assistant", "content": full})
    return

_tool_buf is empty whenever the provider never emitted a __tool_chunk__ — which includes the case where _finish_reason was set to "tool_calls" (core.py:935-936, set from the API's choice.finish_reason in providers/openai_provider.py:92, providers/deepseek_provider.py:208). This is the classic narration-only stall: the model writes "I'll read a.py now…", the provider (or an interrupt mid-retry) closes the stream with tool_calls and an empty array, and koza appends that plan text as the assistant's final message and returns. The task stops half-done, silently, with a plan presented as the answer.

Reproduced against main (core.py unpatched), one narration round then the real call then the answer:

narration-then-call: rounds_started=1 tool_results=[] final_assistant='I will read a.py now.'

One round, no tool ever ran, plan delivered as the answer.

What changed

core.py only (+47/-1):

  • New module constant _DROPPED_TOOLCALL_NUDGE (core.py:427) — the re-prompt text ("Your previous turn indicated a tool call but none was included. Do not narrate a plan or restate intent — issue the actual tool call now…").
  • The if not _tool_buf: branch (core.py:1119-1153) now first checks _finish_reason == "tool_calls". If so and fewer than 3 consecutive stalls have happened, it appends the narration + the nudge (both tagged _dropped_toolcall_nudge) and continues instead of returning. After the budget is spent the narration is delivered as before, so a model that only ever narrates still terminates.
  • _dropped_toolcall_retries is a loop-local (core.py:900), reset whenever a batch actually reaches execution (core.py:1251) and on a genuine text turn (core.py:1151), so the bound counts consecutive stalls.
  • The loop's finally cleanup (core.py:1487) now strips _dropped_toolcall_nudge messages as well as _empty_recovery_synthetic, so the retry scaffolding never lands in the durable transcript.

Provenance

Mechanism ported from the upstream Hermes agent: agent/conversation_loop.py:975-979 (_DROPPED_TOOLCALL_NUDGE_CONTENT) and its consumer agent/turn_final_response.py:279-311 (bounded re-prompt, marked ephemeral scaffolding). Re-implemented in koza's own style — koza name, koza cleanup convention, no imported identifiers.

Verification

Real output from this session (.venv active):

python -m pytest tests -q -p no:cacheprovider   ->  44 passed, 3 warnings in 19.86s
                                                    (main baseline before this change: 38 passed)
ruff check .                                     ->  Found 1991 errors   (identical to main: 1991)
ruff check core.py                               ->  Found 69 errors     (identical to main: 69)
ruff check core.py tests/test_dropped_tool_calls.py -> Found 69 errors  (all 69 in core.py; new test clean)

tests/test_dropped_tool_calls.py (new, force-added — tests/ is in .gitignore): 6 tests, 3 of them driving the real _run_conversation_loop. Test power, core.py rolled back to main via git stash push -- core.py:

narration-then-call: rounds_started=1 tool_results=[] final_assistant='I will read a.py now.'   (main)
narration-then-call: rounds_started=3 tool_results=['read a.py'] final_assistant='Read it.'    (fix)
narration-forever:   rounds_started=1 ...                                                       (main)
narration-forever:   rounds_started=4 ... final_assistant='thinking out loud'                   (fix)

Scratch probe: ~/.hermes/cache/scratch/k04_probe.py (not committed). The bounded-termination case confirms 1 round + at most 3 re-prompts, after which the narration is delivered — no infinite loop (the outer MAX_ROUNDS = 25 also caps it).

Risk / rollback

Low. The change only fires when finish_reason == "tool_calls" and nothing was buffered — a state main currently ends the turn on; every other path is unchanged (a text turn with finish_reason="stop" or None still returns on the first round, covered by two regression tests). Rollback: revert the commit; behaviour returns to main.

Roadmap: K-04

@haydarkadioglu
haydarkadioglu merged commit 514fb4c into main Oct 4, 2026
0 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant