The macOS helper mixes the same two sources into one AAC track (`AudioTrackMixer`) and runs the same discipline, arrived at from the opposite direction. There the emit cursor is anchored to the writer session start and advanced by a monotonic clock read off `CMClockGetHostTimeClock` — the domain ScreenCaptureKit already timestamps its buffers in — with a 10 ms timer on the sample queue moving it while nothing is arriving to move it. The bug it replaces was the anchor, not the counter: `anchor` was set lazily on the first decoded buffer to `max(firstPresentationTime, sessionStart)`, so audio that first arrived four seconds in made frame zero of the track *be* four seconds in. The trailing half was worse still — `drain()` ran only from `ingest` and `finish`, and the helper never called `writer.endSession(atSourceTime:)` at all, so the file ended at the last sample anything happened to deliver. Both halves and any mid-take gap are now one mechanism: chunks go out for as long as the take runs, from whichever source covers them and from silence where none does, and a source is written off for a chunk only once it has gone 250 ms without delivering anything — a decision about when to stop waiting, which shifts nothing, and not a correction, which would. Measured from that source's own last delivery rather than from how far behind the clock its coverage sits, because those differ for a source that is alive but arriving late: a capture path delivering a fixed delay behind the audio it describes would be "behind" forever, so writing it off for that would emit silence and then drop its real samples, every chunk, for the whole take. A pause freezes the clock rather than resetting it, so what follows keeps the position it would have had. Whether ScreenCaptureKit's system-audio output is silence-gapped the way WASAPI loopback is has never been recorded anywhere, and it decides how much of a take the clock is carrying alone rather than how the mixer should work; the helper now reports it per take as `audio-timeline`, whose `undeliveredSeconds` is the whole track if the tap is gapped and near zero if it streams silence. That figure counts what each source handed over rather than what came out of the mixer — holes under two seconds are zero-filled into the source's own buffer, so a hole read off the mixed output would come back covered, which is the exact case it exists to catch — and `droppedSeconds` beside it is the cost side, counting audio that arrived too late to place and reading zero unless the grace is too short for that tap. `swift test` runs the mixer against a collecting sink and a hand-moved clock on every pull request (`swift-macos-helper` in ci.yml, the first Swift job this repository has had); `npm run test:sck-audio-timeline:mac` measures a real tone through a real device, comparing the audio track's length against the video track's in the same file.
0 commit comments