Skip to content

fix(agent): surface the real LLM failure cause and hint an unavailable model - #220

Merged
huhamhire merged 1 commit into
devfrom
fix/llm-error-cause
Sep 7, 2026
Merged

huhamhire merged 1 commit into
devfrom
fix/llm-error-cause

Conversation

@huhamhire

Copy link
Copy Markdown
Owner

Problem

A review failing on the LLM call itself showed up as nothing but 所有备选模型均调用失败 (Failed to generate prediction with any model). The actual cause never reached the run card, because pr-agent''s retry_with_fallback_models logs the underlying exception into loguru''s artifact= field, which the default format does not print.

For the local CLI providers this was compounded twice over:

  • codex reports failures on stdout, not stderr — turn.failed / error events inside its JSONL stream, with stderr left empty. The shim only read stderr on a non-zero exit, so it raised with an empty reason.
  • A turn that dies on a tool/feature failure exits 0 with no assistant message. That empty string flowed downstream into pr-agent''s load_yaml and failed there instead, with no trace of why.

Found while diagnosing a real failure: codex was pinned to a model the account no longer had access to (404 ... The model \x` does not exist or you do not have access to it`), which is precisely the class of failure a user can fix — but only if they are told about it.

Changes

  • shim — extract the cause from the codex event stream on a non-zero exit (precedence turn.failed → last error → error item), via a new optional error_extractor spec key; commands without one keep the stderr fallback.
  • shim — raise explicitly when a zero exit carries an empty reply, naming any error the stream did report, instead of letting it fail downstream as an empty prompt.
  • shim — emit @@MEEBOX_LLM_ERROR@@ {json} on stderr alongside the usage sentinel, so the cause travels past pr-agent''s lossy retry log. Purely additive: the exception still raises and the fallback retry still runs.
  • main — prefer that sentinel over the generic stdout marker for errorMessage, and classify it into a new errorHint (currently only model-unavailable).
  • renderer — render one localized remedy line under the raw cause (four locales). With a local CLI provider the model comes from that CLI''s own configuration, so the fix lives outside the app and the user has to be pointed at it.

Verification

  • lint / typecheck / test / build all pass; four new tests cover the classification, including the negative cases (auth errors, timeouts and empty replies stay unclassified, so no misleading remedy is shown).
  • Exercised end to end against the real codex CLI: a genuine 404 model does not exist run raises with the full cause and emits the sentinel; after switching to an available model the same path succeeds and the usage sentinel is unaffected.

🤖 Generated with Claude Code

…e model

pr-agent's retry_with_fallback_models logs the underlying exception into
loguru's `artifact=` field, which the default format never prints, so a failed
LLM call reached the run card as nothing but "all fallback models failed". For
the local CLI providers this was compounded twice over: codex reports its
failures on stdout (turn.failed / error events in the JSONL stream) and leaves
stderr empty, while a turn that dies on a tool failure exits 0 with no assistant
message at all, which then failed downstream as an unexplained empty prompt.

- shim: extract the cause from the codex event stream on a non-zero exit, and
  raise explicitly when a zero exit carries an empty reply;
- shim: emit `@@MEEBOX_LLM_ERROR@@` on stderr, alongside the usage sentinel, so
  the cause travels past pr-agent's lossy retry log;
- main: prefer that sentinel over the generic marker for errorMessage, and
  classify it into errorHint (currently `model-unavailable`);
- renderer: render one localized remedy line under the raw cause, since a local
  CLI provider's model lives in that CLI's own configuration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@huhamhire huhamhire added bug Something isn't working enhancement New feature or request labels Sep 7, 2026
@huhamhire
huhamhire merged commit a30dc5c into dev Sep 7, 2026
3 checks passed
@huhamhire huhamhire mentioned this pull request Sep 9, 2026
3 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant