Summary
At PreToolUse, Claude Code starts two independent processes — guard-hook.js and collector-hook.js — and both call recordPreToolUse, which read-modify-writes the same pending_spans map in the session's state shard.
The file write itself is atomic (tmp + rename), so no reader ever sees a torn document. The read-modify-write sequence is not: both processes load the same snapshot, each adds its own entry, and whichever saves last silently discards the other's entry.
Measured previously: 3 of 25 entries lost with 50 concurrent processes.
Why it matters
pending_spans records "this tool started, pair it up when it finishes". A lost entry means that tool call's span gets the wrong start time, or fails to pair at all.
Impact is currently limited: tool spans are now emitted eagerly at the post-side event, so the dependency on pending_spans is smaller than it used to be. No guard decision depends on it, and no content is lost.
What it blocks
Tool calls that never reach a span-emitting event — interrupted, or the host killed before PostToolUse — currently have their arguments nowhere but the truncated nio.tool_summary. Closing that gap means emitting spans for whatever is still parked in pending_spans at turn close.
That fix cannot be built on an unreliable map: you cannot distinguish "this entry is gone because the call completed normally and was drained" from "this entry is gone because a concurrent write overwrote it".
Why this is not a quick fix
The obvious shortcut does not apply. The two processes are not writing different fields that could be split into separate files — they call the same function and write the same key.
File locking puts a wait on the guard's critical path, which already carries a 5 s flush budget inside a 10 s hook timeout, and brings lock-timeout and stale-lock handling with it.
Per-process shards need merge logic at every read site (post-side pairing, turn-close reclaim, crash recovery), plus a rule for what happens when the same key appears in both.
The deeper question is why both processes write the same structure at all. guard-hook is recording "the tool I saw when I made a decision"; collector-hook is recording "the tool I am opening a span for". Those are two different facts that happen to share a data structure. Splitting the semantics may be the right fix rather than locking the current shape.
Suggested approach
- Write a test that reliably reproduces the lost update first — without it, any fix is unverifiable, and this repo has caught twenty-one "test passes but verifies nothing" defects.
- Then decide between locking and splitting, with the answer to "should these two facts share a structure" driving it.
Notes
Pre-existing; not introduced by the trace-export work in #9. Deliberately left out of that branch — it touches the state layer shared by all six platforms, which is where four separate paths to "a deny never reaches the host" were found and fixed during that work. Adding an unrelated concurrency change there was not a proportionate risk.
Relevant code: src/scripts/guard-hook.ts (~line 290), src/scripts/lib/collector-core.ts (PreToolUse branch), src/scripts/lib/traces-collector.ts (recordPreToolUse), src/scripts/lib/traces-state-store.ts.
Summary
At
PreToolUse, Claude Code starts two independent processes —guard-hook.jsandcollector-hook.js— and both callrecordPreToolUse, which read-modify-writes the samepending_spansmap in the session's state shard.The file write itself is atomic (tmp + rename), so no reader ever sees a torn document. The read-modify-write sequence is not: both processes load the same snapshot, each adds its own entry, and whichever saves last silently discards the other's entry.
Measured previously: 3 of 25 entries lost with 50 concurrent processes.
Why it matters
pending_spansrecords "this tool started, pair it up when it finishes". A lost entry means that tool call's span gets the wrong start time, or fails to pair at all.Impact is currently limited: tool spans are now emitted eagerly at the post-side event, so the dependency on
pending_spansis smaller than it used to be. No guard decision depends on it, and no content is lost.What it blocks
Tool calls that never reach a span-emitting event — interrupted, or the host killed before
PostToolUse— currently have their arguments nowhere but the truncatednio.tool_summary. Closing that gap means emitting spans for whatever is still parked inpending_spansat turn close.That fix cannot be built on an unreliable map: you cannot distinguish "this entry is gone because the call completed normally and was drained" from "this entry is gone because a concurrent write overwrote it".
Why this is not a quick fix
The obvious shortcut does not apply. The two processes are not writing different fields that could be split into separate files — they call the same function and write the same key.
File locking puts a wait on the guard's critical path, which already carries a 5 s flush budget inside a 10 s hook timeout, and brings lock-timeout and stale-lock handling with it.
Per-process shards need merge logic at every read site (post-side pairing, turn-close reclaim, crash recovery), plus a rule for what happens when the same key appears in both.
The deeper question is why both processes write the same structure at all.
guard-hookis recording "the tool I saw when I made a decision";collector-hookis recording "the tool I am opening a span for". Those are two different facts that happen to share a data structure. Splitting the semantics may be the right fix rather than locking the current shape.Suggested approach
Notes
Pre-existing; not introduced by the trace-export work in #9. Deliberately left out of that branch — it touches the state layer shared by all six platforms, which is where four separate paths to "a deny never reaches the host" were found and fixed during that work. Adding an unrelated concurrency change there was not a proportionate risk.
Relevant code:
src/scripts/guard-hook.ts(~line 290),src/scripts/lib/collector-core.ts(PreToolUse branch),src/scripts/lib/traces-collector.ts(recordPreToolUse),src/scripts/lib/traces-state-store.ts.