Repository navigation
[Bug] zcode CLI TUI: unbounded native memory growth during model streaming — long sessions reach tens of GB #913
Copy link
Copy link
Open
Labels
priority: P2普通优先级普通优先级
Description
Activity
👋 感谢你的反馈,我们已经收到。
- 维护者看到后会尽快回复你。
- 状态保持为
status: 待评估,你可以随时补充信息。 - 信息不全时我们会打上
needs: 更多信息标签并 @ 你。
👋 Thanks — we've received your issue.
- A maintainer will get back to you as soon as we can.
- Status stays at
status: 待评估(Triage); feel free to add context. - If we need more details, we'll add
needs: 更多信息and ping you.
Follow-up with the exact workaround we're running on 3.14.3 (numbers in the report above). Three small changes against the TUI source layout; feel free to adapt.
1. Delta coalescing (new module,
app-model-streaming.ts):const STREAM_FLUSH_MS = 500; type StreamDeltaSetters = { setLiveModelText: (update: (current: string) => string) => void; setMessages: (update: (messages: Message[]) => Message[]) => void; setStatus: (status: string) => void; }; const pendingTextById = new Map<string, string>(); const pendingThoughtById = new Map<string, string>(); let pendingLiveText = ""; let flushTimer: ReturnType<typeof setTimeout> | undefined; let flushSetters: StreamDeltaSetters | undefined; /** * Buffers model_streaming text/reasoning deltas; returns true when the event * was consumed. The buffer is applied on a STREAM_FLUSH_MS timer, so React * re-renders and native opentui commits happen a few times per second instead * of once per delta. */ export function coalesceModelStreamingDelta( payload: Record<string, unknown>, handlers: StreamDeltaSetters, ): boolean { const kind = stringField(payload, "kind"); const delta = stringField(payload, "delta") ?? ""; let status: string; if (kind === "reasoning_delta") { const id = stringField(payload, "assistantMessageId"); if (!id) return false; pendingThoughtById.set(id, (pendingThoughtById.get(id) ?? "") + delta); status = "Streaming model reasoning..."; } else if (kind === "text_delta" || (!kind && delta)) { const id = stringField(payload, "assistantMessageId"); if (id) pendingTextById.set(id, (pendingTextById.get(id) ?? "") + delta); else pendingLiveText += delta; status = "Streaming model response..."; } else { return false; } flushSetters = handlers; handlers.setStatus(status); // same-string setState is a React no-op if (flushTimer === undefined) { flushTimer = setTimeout(() => { flushTimer = undefined; flushModelStreamingDeltas(); }, STREAM_FLUSH_MS); (flushTimer as { unref?: () => void }).unref?.(); } return true; } export function flushModelStreamingDeltas(): void { if (flushTimer !== undefined) { clearTimeout(flushTimer); flushTimer = undefined; } const setters = flushSetters; if (setters === undefined) return; flushSetters = undefined; if (pendingLiveText) { const live = pendingLiveText; pendingLiveText = ""; setters.setLiveModelText((current) => `${current}${live}`); } for (const [id, text] of pendingTextById) { pendingTextById.delete(id); setters.setMessages((messages) => appendStreamingTextDelta(messages, id, text)); } for (const [id, thought] of pendingThoughtById) { pendingThoughtById.delete(id); setters.setMessages((messages) => appendStreamingThoughtDelta(messages, id, thought)); } }
2. Event-pipeline hook (in the session-event applier, after dedupe):
if (!rememberSessionEvent(applied, event)) return false; // Streaming deltas go to the coalescing buffer (after dedupe, so duplicate // delivery is still blocked by event id); every other event flushes the buffer // first, keeping tool/thought/assistant rows ordered relative to streamed text. if (event.type === "model_streaming" && coalesceModelStreamingDelta(asRecord(event.payload), input)) { input.setLastEvent(describeSessionEvent(event)); // constant string -> no-op render return true; } flushModelStreamingDeltas(); input.setLastEvent(describeSessionEvent(event)); applySessionEventToState(event, input, input.copy); return true;
3. Memoized transcript rows:
function MessageRowImpl(props: MessageRowProps) { /* unchanged body */ } // appendStreamingTextDelta / appendStreamingThoughtDelta are immutable maps that // preserve the object identity of untouched messages, so memo skips every // settled row and only the actively streaming row re-renders per flush. export const MessageRow = React.memo(MessageRowImpl);
Two extra observations that may help whoever digs into the native side:
- The delta-append path (
appendTextPart/appendThoughtPart) already extends the last part of the target message instead of creating a part per delta — part count stays constant per message, which is what makes the memo effective. - We also tested rendering plain text instead of markdown while streaming: it did not reduce retention (within run-to-run noise), so the per-flush cost appears to sit in the commit/marshalling path itself rather than in markdown parsing.
- Ordering is preserved because non-streaming events flush the buffer before being applied; the flush timer is
unref()'d so it never holds the process open, and at mostSTREAM_FLUSH_MSof projected text can be lost if the TUI crashes mid-turn (the authoritative transcript rebuilds it from the session store).
- The delta-append path (
Metadata
Metadata
Assignees
Labels
priority: P2普通优先级普通优先级
Environment
Symptom
A long-lived interactive CLI session grows RSS continuously with usage. Worst observed case: ~12 h session reached 27.4 GB RSS + 31.4 GB swap (VmPeak 121 GB) — while the session's entire transcript in
db.sqliteis only 5 MB (~12,000× amplification). An idle TUI is completely flat (~455 MB over 9 min), so growth is driven by turns, not time. Nothing external reclaims it: no GC pressure signal,malloc_trimfrees ~nothing, and the pages are never returned. Only exiting andzcode --resumerecovers.Reproduction
rw-pmappings of ~0.5 MB each (visible in/proc/<pid>/maps; hundreds of new regions per heavy turn).malloc_trimno-op), the worker thread, or the transcript itself.Diagnosis
Every streamed delta (
model_streaming→text_delta/reasoning_delta) triggers a full-transcript React re-render plus a native commit through @mbears/opentui —MessageRowis not memoized, so the whole transcript reconciles per delta, and the streaming message's full content is re-marshalled through koffi intolibopentui.soon every commit. Retention scales with flush-count × document size.A headless control (
zcode -p, same essay, no TUI — confirmed the headless process does not even maplibopentui.so) still retains ~+70 MB per heavy turn, so there is a second, smaller accumulation term inside the agent streaming core, independent of rendering.Heavy-turn retention also exhibits a deferred release: memory peaks mid-turn (~+110–170 MB) and a chunk of the native regions is released some time after, so single-shot measurements vary widely (+31…+284 MB observed for the same workload).
Workaround we shipped locally (3.14.3,
@zcode/tuibundle)unref()'d.MessageRow = React.memo(MessageRowImpl)(the delta-append path already preserves object identity of untouched messages, so settled rows stop re-rendering).Result: tiny-turn retention +33 → ~+8.5 MB/turn RSS (per-turn map growth 127 → 42; repeat-turn 132 → 8,
94%); heavy essay turns drop to **+33 / +0 / +73 MB over three consecutive turns (+38 MB/turn net average)** vs up to +284 MB unpatched.Ask
Please track down the native retention in the opentui commit path (full-content re-marshalling per update) and the per-byte retention in the agent streaming core. Delta coalescing + row memoization would be a cheap, significant improvement to fold into the shipped TUI. Happy to share the pty repro driver and per-15s RSS/maps sampler if useful.