Summary
Community observation and verification by Discord user Dillon:
"I've a local AI gateway to inspect token usage for each LLM call. And I found when using ZCode, and with GLM models, there is a huge waste of tokens... in thinking mode, the thinking result text is reinjected into all following calls... Similar, the tool response is also reinjected into all following calls. I've vibe coded a local proxy to help filtering out these 2 cases of wasted tokens, saving 30% + 25% tokens!"
Problem Analysis
When running multi-turn agent sessions in ZCode with reasoning/thinking models:
- Re-injection of Historical
reasoning_content: The thinking scratchpad from prior assistant turns is accumulated and resubmitted in every subsequent API turn. In standard LLM agent harness conventions, thinking traces are generation-time scratchpads and are pruned from conversational history once the assistant produces its final response or tool call.
- Accumulation of Obsolete Tool Output Dumps: Full outputs from tools executed many turns ago (large file reads, verbose shell outputs) remain unpruned in the context array, compounding token growth on every subsequent turn.
Measured Empirical Data (from Dillon's proxy logs)
Using an interception proxy (reasoning-strip-proxy) between ZCode and the gateway:
- Reasoning token savings: ~30% reduction by stripping prior
reasoning_content blocks (stripped=68 (~284k chars) per request cycle).
- Tool output pruning savings: ~25% reduction by pruning obsolete historical tool messages (
pruned=8 (~39k chars)).
- Combined token economy: Up to 50%+ reduction in total prompt tokens across long multi-turn sessions without degrading completion quality.
Proposed Architectural Optimization
Bake this pruning directly into ZCode's native conversation serializer before dispatching /v1/chat/completions (or equivalent client requests):
- Strip Historical Reasoning: For all previous assistant turns (turn index < current turn), exclude the
reasoning_content / thinking block from the serialized request payload.
- Context-Aware Tool Pruning: Truncate or prune verbose tool outputs from turns older than
N steps, keeping only essential summaries or the latest relevant tool returns.
User Impact & Benefits
- Cuts prompt token consumption by up to 50% on long agent tasks.
- Prevents users from burning through their 5-hour quota and weekly limits prematurely.
- Significantly lowers request latency by reducing prompt prefill overhead on subsequent calls.
Summary
Community observation and verification by Discord user Dillon:
Problem Analysis
When running multi-turn agent sessions in ZCode with reasoning/thinking models:
reasoning_content: The thinking scratchpad from prior assistant turns is accumulated and resubmitted in every subsequent API turn. In standard LLM agent harness conventions, thinking traces are generation-time scratchpads and are pruned from conversational history once the assistant produces its final response or tool call.Measured Empirical Data (from Dillon's proxy logs)
Using an interception proxy (
reasoning-strip-proxy) between ZCode and the gateway:reasoning_contentblocks (stripped=68 (~284k chars)per request cycle).pruned=8 (~39k chars)).Proposed Architectural Optimization
Bake this pruning directly into ZCode's native conversation serializer before dispatching
/v1/chat/completions(or equivalent client requests):reasoning_content/ thinking block from the serialized request payload.Nsteps, keeping only essential summaries or the latest relevant tool returns.User Impact & Benefits