Skip to content

[Bug] zcode CLI TUI: unbounded native memory growth during model streaming — long sessions reach tens of GB #913

Description

@rohitnanda1443

Environment

  • zcode-cli 3.14.3 (system install, bundled Node 26), Gentoo Linux x64, terminal TUI (also observed under a pty driver)
  • Host: 60 GB RAM + 57 GB swap

Symptom

A long-lived interactive CLI session grows RSS continuously with usage. Worst observed case: ~12 h session reached 27.4 GB RSS + 31.4 GB swap (VmPeak 121 GB) — while the session's entire transcript in db.sqlite is only 5 MB (~12,000× amplification). An idle TUI is completely flat (~455 MB over 9 min), so growth is driven by turns, not time. Nothing external reclaims it: no GC pressure signal, malloc_trim frees ~nothing, and the pages are never returned. Only exiting and zcode --resume recovers.

Reproduction

  1. Tiny turns: drive the TUI over a pty with 30 × "reply OK" turns → ~5–7 MB retained per turn, linear, no plateau (554→627 MB across turns 12–23).
  2. Heavy turns: a single 2500-word essay prompt → up to +284 MB native retention in one turn. The growth is in anonymous private rw-p mappings of ~0.5 MB each (visible in /proc/<pid>/maps; hundreds of new regions per heavy turn).
  3. What it is NOT: V8 JS heap (heap snapshots across turn counts show flat retained JS, ~10 MB), glibc arenas (malloc_trim no-op), the worker thread, or the transcript itself.

Diagnosis

Every streamed delta (model_streaming → text_delta / reasoning_delta) triggers a full-transcript React re-render plus a native commit through @mbears/opentui — MessageRow is not memoized, so the whole transcript reconciles per delta, and the streaming message's full content is re-marshalled through koffi into libopentui.so on every commit. Retention scales with flush-count × document size.

A headless control (zcode -p, same essay, no TUI — confirmed the headless process does not even map libopentui.so) still retains ~+70 MB per heavy turn, so there is a second, smaller accumulation term inside the agent streaming core, independent of rendering.

Heavy-turn retention also exhibits a deferred release: memory peaks mid-turn (~+110–170 MB) and a chunk of the native regions is released some time after, so single-shot measurements vary widely (+31…+284 MB observed for the same workload).

Workaround we shipped locally (3.14.3, @zcode/tui bundle)

  1. Coalesce streaming deltas: buffer text/reasoning deltas and apply them on a 500 ms timer; any non-streaming event flushes the buffer first, preserving tool/thought/text ordering; timer unref()'d.
  2. MessageRow = React.memo(MessageRowImpl) (the delta-append path already preserves object identity of untouched messages, so settled rows stop re-rendering).

Result: tiny-turn retention +33 → ~+8.5 MB/turn RSS (per-turn map growth 127 → 42; repeat-turn 132 → 8, 94%); heavy essay turns drop to **+33 / +0 / +73 MB over three consecutive turns (+38 MB/turn net average)** vs up to +284 MB unpatched.

Ask

Please track down the native retention in the opentui commit path (full-content re-marshalling per update) and the per-byte retention in the agent streaming core. Delta coalescing + row memoization would be a cheap, significant improvement to fold into the shipped TUI. Happy to share the pty repro driver and per-15s RSS/maps sampler if useful.

Activity

  1. github-actions commented on Oct 3, 2026

    @github-actions

    👋 感谢你的反馈,我们已经收到。

    • 维护者看到后会尽快回复你。
    • 状态保持为 status: 待评估,你可以随时补充信息。
    • 信息不全时我们会打上 needs: 更多信息 标签并 @ 你。

    👋 Thanks — we've received your issue.

    • A maintainer will get back to you as soon as we can.
    • Status stays at status: 待评估 (Triage); feel free to add context.
    • If we need more details, we'll add needs: 更多信息 and ping you.
  2. rohitnanda1443 commented on Oct 3, 2026

    @rohitnanda1443
    Author

    Follow-up with the exact workaround we're running on 3.14.3 (numbers in the report above). Three small changes against the TUI source layout; feel free to adapt.

    1. Delta coalescing (new module, app-model-streaming.ts):

    const STREAM_FLUSH_MS = 500;
    
    type StreamDeltaSetters = {
      setLiveModelText: (update: (current: string) => string) => void;
      setMessages: (update: (messages: Message[]) => Message[]) => void;
      setStatus: (status: string) => void;
    };
    
    const pendingTextById = new Map<string, string>();
    const pendingThoughtById = new Map<string, string>();
    let pendingLiveText = "";
    let flushTimer: ReturnType<typeof setTimeout> | undefined;
    let flushSetters: StreamDeltaSetters | undefined;
    
    /**
     * Buffers model_streaming text/reasoning deltas; returns true when the event
     * was consumed. The buffer is applied on a STREAM_FLUSH_MS timer, so React
     * re-renders and native opentui commits happen a few times per second instead
     * of once per delta.
     */
    export function coalesceModelStreamingDelta(
      payload: Record<string, unknown>,
      handlers: StreamDeltaSetters,
    ): boolean {
      const kind = stringField(payload, "kind");
      const delta = stringField(payload, "delta") ?? "";
      let status: string;
      if (kind === "reasoning_delta") {
        const id = stringField(payload, "assistantMessageId");
        if (!id) return false;
        pendingThoughtById.set(id, (pendingThoughtById.get(id) ?? "") + delta);
        status = "Streaming model reasoning...";
      } else if (kind === "text_delta" || (!kind && delta)) {
        const id = stringField(payload, "assistantMessageId");
        if (id) pendingTextById.set(id, (pendingTextById.get(id) ?? "") + delta);
        else pendingLiveText += delta;
        status = "Streaming model response...";
      } else {
        return false;
      }
      flushSetters = handlers;
      handlers.setStatus(status); // same-string setState is a React no-op
      if (flushTimer === undefined) {
        flushTimer = setTimeout(() => {
          flushTimer = undefined;
          flushModelStreamingDeltas();
        }, STREAM_FLUSH_MS);
        (flushTimer as { unref?: () => void }).unref?.();
      }
      return true;
    }
    
    export function flushModelStreamingDeltas(): void {
      if (flushTimer !== undefined) {
        clearTimeout(flushTimer);
        flushTimer = undefined;
      }
      const setters = flushSetters;
      if (setters === undefined) return;
      flushSetters = undefined;
      if (pendingLiveText) {
        const live = pendingLiveText;
        pendingLiveText = "";
        setters.setLiveModelText((current) => `${current}${live}`);
      }
      for (const [id, text] of pendingTextById) {
        pendingTextById.delete(id);
        setters.setMessages((messages) => appendStreamingTextDelta(messages, id, text));
      }
      for (const [id, thought] of pendingThoughtById) {
        pendingThoughtById.delete(id);
        setters.setMessages((messages) => appendStreamingThoughtDelta(messages, id, thought));
      }
    }

    2. Event-pipeline hook (in the session-event applier, after dedupe):

    if (!rememberSessionEvent(applied, event)) return false;
    // Streaming deltas go to the coalescing buffer (after dedupe, so duplicate
    // delivery is still blocked by event id); every other event flushes the buffer
    // first, keeping tool/thought/assistant rows ordered relative to streamed text.
    if (event.type === "model_streaming" && coalesceModelStreamingDelta(asRecord(event.payload), input)) {
      input.setLastEvent(describeSessionEvent(event)); // constant string -> no-op render
      return true;
    }
    flushModelStreamingDeltas();
    input.setLastEvent(describeSessionEvent(event));
    applySessionEventToState(event, input, input.copy);
    return true;

    3. Memoized transcript rows:

    function MessageRowImpl(props: MessageRowProps) {
      /* unchanged body */
    }
    // appendStreamingTextDelta / appendStreamingThoughtDelta are immutable maps that
    // preserve the object identity of untouched messages, so memo skips every
    // settled row and only the actively streaming row re-renders per flush.
    export const MessageRow = React.memo(MessageRowImpl);

    Two extra observations that may help whoever digs into the native side:

    • The delta-append path (appendTextPart/appendThoughtPart) already extends the last part of the target message instead of creating a part per delta — part count stays constant per message, which is what makes the memo effective.
    • We also tested rendering plain text instead of markdown while streaming: it did not reduce retention (within run-to-run noise), so the per-flush cost appears to sit in the commit/marshalling path itself rather than in markdown parsing.
    • Ordering is preserved because non-streaming events flush the buffer before being applied; the flush timer is unref()'d so it never holds the process open, and at most STREAM_FLUSH_MS of projected text can be lost if the TUI crashes mid-turn (the authoritative transcript rebuilds it from the session store).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions