Skip to content

[Feature / Optimization] ZCode Runtime: Strip historical reasoning traces & prune stale tool results across multi-turn sessions (cuts ~50% token bloat) #925

Description

@romangalaxys10-spec

Summary

Community observation and verification by Discord user Dillon:

"I've a local AI gateway to inspect token usage for each LLM call. And I found when using ZCode, and with GLM models, there is a huge waste of tokens... in thinking mode, the thinking result text is reinjected into all following calls... Similar, the tool response is also reinjected into all following calls. I've vibe coded a local proxy to help filtering out these 2 cases of wasted tokens, saving 30% + 25% tokens!"

Problem Analysis

When running multi-turn agent sessions in ZCode with reasoning/thinking models:

  1. Re-injection of Historical reasoning_content: The thinking scratchpad from prior assistant turns is accumulated and resubmitted in every subsequent API turn. In standard LLM agent harness conventions, thinking traces are generation-time scratchpads and are pruned from conversational history once the assistant produces its final response or tool call.
  2. Accumulation of Obsolete Tool Output Dumps: Full outputs from tools executed many turns ago (large file reads, verbose shell outputs) remain unpruned in the context array, compounding token growth on every subsequent turn.

Measured Empirical Data (from Dillon's proxy logs)

Using an interception proxy (reasoning-strip-proxy) between ZCode and the gateway:

  • Reasoning token savings: ~30% reduction by stripping prior reasoning_content blocks (stripped=68 (~284k chars) per request cycle).
  • Tool output pruning savings: ~25% reduction by pruning obsolete historical tool messages (pruned=8 (~39k chars)).
  • Combined token economy: Up to 50%+ reduction in total prompt tokens across long multi-turn sessions without degrading completion quality.

Proposed Architectural Optimization

Bake this pruning directly into ZCode's native conversation serializer before dispatching /v1/chat/completions (or equivalent client requests):

  1. Strip Historical Reasoning: For all previous assistant turns (turn index < current turn), exclude the reasoning_content / thinking block from the serialized request payload.
  2. Context-Aware Tool Pruning: Truncate or prune verbose tool outputs from turns older than N steps, keeping only essential summaries or the latest relevant tool returns.

User Impact & Benefits

  • Cuts prompt token consumption by up to 50% on long agent tasks.
  • Prevents users from burning through their 5-hour quota and weekly limits prematurely.
  • Significantly lowers request latency by reducing prompt prefill overhead on subsequent calls.

Activity

  1. github-actions commented on Oct 5, 2026

    @github-actions

    👋 感谢你的反馈,我们已经收到。

    • 维护者看到后会尽快回复你。
    • 状态保持为 status: 待评估,你可以随时补充信息。
    • 信息不全时我们会打上 needs: 更多信息 标签并 @ 你。

    👋 Thanks — we've received your issue.

    • A maintainer will get back to you as soon as we can.
    • Status stays at status: 待评估 (Triage); feel free to add context.
    • If we need more details, we'll add needs: 更多信息 and ping you.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions