Skip to content

fix(core): refuse tool calls with unparseable JSON arguments instead of executing them [K-02] - #5

Merged
haydarkadioglu merged 1 commit into
mainfrom
feat/K-02-invalid-tool-json
Oct 1, 2026
Merged

haydarkadioglu merged 1 commit into
mainfrom
feat/K-02-invalid-tool-json

Conversation

@haydarkadioglu

Copy link
Copy Markdown
Owner

Summary

A tool call whose argument JSON cannot be parsed is now refused instead of executed with empty
arguments
, and the parse error is fed back to the model so it can re-issue the call. Repeating
broken JSON ends the turn instead of looping.

Why

core.py:1099-1112 built each call's argument dict like this:

raw_args = stc["args"] or "{}"
try:
    args_parsed = _json.loads(raw_args)
except Exception:
    repaired = _repair_tool_call_arguments(raw_args, tool_name=stc["name"])
    try:
        args_parsed = _json.loads(repaired)
    except Exception:
        args_parsed = {}          # <-- silent degradation

_repair_tool_call_arguments returns its failure sentinel "{}" (utils/json_utils.py:96) when it
cannot repair a payload, so the inner json.loads succeeds and the tool runs on {} — with
every parameter missing. A truncated payload is the worst case: repair closes the braces it was cut
before, so half-streamed arguments become valid-looking JSON.

Measured with a scratch probe against the real _run_conversation_loop (mechanism emulated so the
old path is callable, see Verification):

NEW  executed: []                                   # after this change
OLD  executed: [('run_terminal', {})]               # before: empty args reached the tool

The user-visible symptom is a "tool bug" — a tool acting on nothing — while the real cause (broken
JSON from the model) stays invisible. There was also no truncation refusal: args cut by the output
limit or a dropped stream were indistinguishable from a parameterless call.

What changed

  • utils/json_utils.py — new parse_tool_call_arguments(raw_args, tool_name) returning
    (arguments, error): error is None only when the payload parsed (after the existing repair
    passes). A payload that does not close its JSON object/array is reported as truncated before
    repair runs (tool_call_arguments_look_truncated); non-object payloads are rejected too. Empty
    payloads stay a legal parameterless call.
  • core.py — the loop calls the new parser. If any call in a round has unparseable arguments,
    nothing in that round executes: the invalid call gets a tool result carrying the reason, and
    its siblings get an explicit "Skipped" result, so role/tool-result pairing stays intact. The model
    then gets another round to re-issue the call, bounded at 3 consecutive rounds; past that the turn
    ends with a visible notice instead of looping to MAX_ROUNDS.
  • core.py — removed the now-unused json as _json local import and the two stale utils.json_utils
    imports, replacing them with the one symbol actually used.

Provenance

Ported (mechanism and edge cases, re-implemented in koza's style) from
hermes:agent/turn_tool_validation.py:148-223 — invalid/truncated argument validation, the
truncation refusal, and recovery results used to re-prompt the model.

Verification

. .venv/bin/activate
python -m pytest tests -q -p no:cacheprovider
  -> 29 passed, 3 warnings in 19.82s        (baseline on main before this change: 19 passed)
ruff check .
  -> Found 1991 errors                     (baseline on main: 1995 — pre-existing repo-wide)
ruff check core.py utils/json_utils.py
  -> Found 78 errors                       (baseline: 82)
ruff check tests/test_tool_call_arguments.py
  -> All checks passed!

No test that was red on main became red here; the 10 new tests in
tests/test_tool_call_arguments.py (force-added — tests/ is gitignored) cover: empty/repairable/
truncated/non-object payloads, a truncated call never reaching the tool, a valid sibling being
skipped in a broken round, the bounded stop after repeated broken JSON, and repair still executing
normally. Four of them drive the real _run_conversation_loop with a stubbed streaming provider.

Live probe (_run_conversation_loop, truncated payload {"command": "rm -rf /tmp):

Invalid tool-call arguments (round 1/3), executing nothing: truncated JSON — the payload was cut off
before closing its object: '{"command": "rm -rf /tmp'
NEW  executed: []
NEW  tool result: Error: invalid JSON in arguments for 'run_terminal': truncated JSON — ...
OLD  executed: [('run_terminal', {})]

Risk / rollback

Contained to the tool-call argument path in core.py and two new pure functions in
utils/json_utils.py; no dependency, config or prompt change. Behavioural change is intentional:
a broken-JSON call no longer executes — it is reported to the model and retried, then the turn stops.
Rollback: revert this commit; _repair_tool_call_arguments is untouched and still exported.

Roadmap

Roadmap: K-02

@haydarkadioglu
haydarkadioglu merged commit f376e54 into main Oct 1, 2026
0 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant