You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat(debates): time-align claims and surface them over the video as they are said (GEO-2909) - #2439
Claims are a debate's payload and they lived behind a button, in a side panel, ordered by ranking score, with no relation to what was on screen. This puts them where the attention is.
The end-of-debate card is not in this PR — it is #2459, stacked on this branch, so the two can be judged apart. Recombining is one git revert.
What you get
Claims appear over the video once they have been said. Each debater has their own corner in the bottom-right of their own tile, built to the Figma frame: a glass card carrying their avatar and name, the crowd's share, the claim clamped to two lines, and two small icons — thumbs for a stance claim, chevrons for a factual one.
One corner per debater rather than one for the player, because a viewer is looking at whoever is talking; a shared corner asks the eye to leave the speaker in order to read what the speaker is saying.
A card appears at the claim's end, not its start. Anchored to the start it arrived while the debater was still mid-sentence — asserting a claim the viewer had not heard them make yet, and asking them to agree with something still being argued.
The corner rests on a chip, and opens into the backlog. Cards expire, so the corner is empty for most of a debate; a small "N claims" pill says the backlog is there. Pointing at that debater's tile, tabbing into the stack, or pressing the chip opens everything they have said so far — scrollable, newest at the bottom, dissolving into the tile's edge. Bounded by the playhead, so it is a record rather than a table of contents.
The chip exists because hover is nothing on a touch screen and close to nothing on a desktop — an invisible affordance is only found by accident. On touch the chip is deliberately the only way in: the tile's hover handler ignores anything but a real mouse, or the synthesised pointerenter from a tap would throw the corner open every time someone tapped to pause.
Markers on the scrubber, one per precisely-placed claim, each a seek — at the claim's end, where its card appears.
The claims panel is in spoken order. It sorted by ranking score, which is right for a feed of unrelated claims and wrong inside a transcript: a debate is an argument, and reading it out of sequence loses the thread. Ties — every claim of one turn — still fall to the ranking. Rows whose moment is firm enough carry a timecode.
How it knows when a claim was said
core/debates/claim-timing.ts — one resolver, three sources in order:
Source
Where from
Accuracy
published
timecodes on the block → claim relation entity
exact
segment
the claim's content words matched against the Whisper transcript
right span, loose at the edges, carries a confidence
block
the turn's own window
coarse, never wrong
So a debate published long before any of this still gets an answer, and a better one the day offsets land for it. Both reads are shared cache entries, not new requests: claims come back on the key the count badge and panel already use, the transcript on the key the player already fetches for subtitles.
One bar decides whether a moment may be stated.isAssertableMoment — a timecode and a card over the video both go through it. They used to disagree: the live layer asked for 0.55 while the panel's timecode asked only that the claim had been matched at all, so a 0.40 match judged too loose to draw over the video still printed a time to the second, indistinguishable from a published offset. Ordering deliberately does not go through it: a list has to put a claim somewhere, and where it lands asserts far less than a timestamp.
Four findings worth a look
A block's markdown is the verbatim concatenation of its Whisper segments, so locating a turn is string equality, not fuzzy matching. All six blocks matched exactly. That is why the block fallback is trustworthy rather than a guess.
Relation position is not transcript order.debate-publish-draft.ts assigns Position.generate(), which is random rather than monotonic. On the test debate the blocks in position order are the 3rd, 1st, 6th, 4th, 5th and 2nd turns — so the ordering transcript-claims.ts documented as "transcript order" was arbitrary on every debate since it shipped. Fixed, and the panel now orders by recovered time.
The scrubber markers were misplaced. Their fractions were scaled by the latest claim's end while the track is sized by the debate's timeline. Those agree only when the last claim runs to the final second — true of the hand-published test debate and false elsewhere. Across eight live debates the gap reached 11%: about 20 seconds out on a 210-second debate, on a control whose whole job is landing you somewhere exact.
scoreWindow could exceed 1. A repeated claim word was counted once per occurrence against a set of window words, so its precision term could pass 1 and push the score past the [0, 1] it documents. Two claims scored 1.05 and 1.03; a claim saying "AI" three times outranked a tighter match saying it once.
Verified against live data
bun scripts/verify-claim-timing.ts <debateEntityId> <spaceId> runs the real grouping, resolver and ticker selection over testnet. On the test debate all 13 claims resolve published; on one with none, the three tiers come apart as they should:
claims: 9 with published timecodes: 0
by source: block=4 segment=5
firm enough to state — timecode, live card: 3 of 9
used for ordering only (matched, below the bar): 2
Behaviour was checked in a real browser rather than only in jsdom — which is how the line-clamp override, the stale ResizeObserver, and the top-anchored open list were all caught.
Backfill
scripts/plan-claim-timecodes.ts produced the plan the publishing agent ran: 140 writes across 50 debates at a 0.7 confidence floor. The floor is the point — published is the resolver's most-trusted source, scored 1.0, so writing a weak match launders a guess into an exact answer no later gate can demote.
This has since been run: the graph now holds 153 published timecodes (13 hand-published + 140). Re-running the planner produces an empty plan, which is the idempotency working.
The planning documents and the generated plan JSON were working papers for that one-off publish and are not committed — Preston asked for them out of the PR. The scripts are self-describing and stay; the record of which offsets were written, and at what confidence, lives with the run rather than in the repo. Worth knowing for GEO-2958, which will want to know which offsets it may overwrite with the extractor's real spans.
Test data
Timecodes for the test debate were hand-published first so the published path had something real to read — thirteen claims, four of them corrected by hand after reading the transcript.
Reusing the existing Start offset / End offset Integer properties rather than minting new ones was Preston's call and a good one — they belong to the Selector type the graph already uses to anchor a Reply to at a span of its target. Note the API serialises Integer values as strings; parsed and range-checked with a test per failure mode, because a half-parsed pair would be drawn as a real moment.
Testing
claim-timing.test.ts — a verbatim turn of the real debate; each of its three claims must land on the sentence that says it. Real data, because the thing under test is whether an extractor's paraphrase matches back to speech, and a hand-written fixture would quietly be written to match. Plus the scorer's range, the assertable bar, and spoken-order sorting.
claim-ticker.test.ts — the end-anchored window, stack capping, expiry, the backlog, fades, and marker placement against the timeline.
debate-claim-ticker.test.tsx — attribution on the card, the vocabulary verb on the share, the expand toggle, the chip, and no play/pause leak.
debate-claims-panel.test.tsx — spoken order beating the ranking, and the ranking still breaking ties.
transcript-claims.test.ts — offset parsing, six discard cases.
Full apps/web suite: 483 files, 5708 tests, green.tsc --noEmit clean apart from the four files already failing on master.
⚠️ The suite is flaky under full parallelism in a way unrelated to this branch — a different handful of files fails on each full run with "Tests must not reach the network", and all of them pass in isolation. Worth its own ticket; it will make CI here look red at random.
Self-review pass
e7d1f87 removes the duplication this PR introduced: three copies of the claims query collapsed to one shared by the app and both scripts, a hand-rolled traversal in the planner replaced by the relation entity id the grouping now carries, and the useClaimResponseState + useClaimPositionControl wiring — repeated at five call sites, two of them added here — extracted into useDebateClaimResponse. Net −224 lines, no behaviour change.
Not in this PR
Nothing publishes offsets automatically yet. GEO-2958 covers geo-chat emitting start_ms/end_ms per extracted claim, with the rule that where the extractor knows the span it should report it rather than leave it to be guessed.
Running it locally
The debates feed needs geo-chat, which apps/web/.env does not point anywhere by default:
jwalkingjew
changed the title
feat(debates): resolve when each claim was said, so claims can be tied to the video (GEO-2909)
feat(debates): surface claims over the video as they are said, with Agree/Disagree (GEO-2909)
Sep 15, 2026
jwalkingjew
changed the title
feat(debates): surface claims over the video as they are said, with Agree/Disagree (GEO-2909)
feat(debates): time-align claims and ask for positions when the debate ends (GEO-2909)
Sep 15, 2026
jwalkingjew
changed the title
feat(debates): time-align claims and ask for positions when the debate ends (GEO-2909)
feat(debates): time-align claims, surface them during playback, and score the debate at the end (GEO-2909)
Sep 16, 2026
d63b9806b — the "Debates" frame (76391:24940) in Geo — Web, pulled through the Figma Dev Mode MCP server so the fills, spacing and type are the frame's own values rather than eyeballed off a screenshot.
This replaces the chyron-line treatment of the live claim layer and reworks the chrome around it. Both hosts get it: the explore feed card and the fullscreen debates feed render the same DebateFeedPlayer.
Claim cards — one stack in the player's bottom-left, whoever is speaking, instead of one per debater's corner. Each card carries the speaker's avatar and name, the crowd's share, and the two answer icons on its top line; the claim is clamped to three lines below. The card attributing itself is what frees the stack to have a fixed home rather than jumping between halves mid-sentence.
Chrome — tiles flush with one 12px outer radius, so the subtitle can straddle the seam (the one strip that is never a face, and the one place it cannot collide with the claim stack). Play/pause and mute as two 40px circles top-left, up when the debate is idle and hover-only while it runs. Countdown badge gets the frame's stroked ring. Debater name moves bottom-right on a scrim that runs to black.
Two displacements, confirmed rather than guessed. The Winner? pill and the position chip are gone from the tile — the name row took that corner, and winner voting already lives on the end-of-debate scorecard and in the claims panel (WinnerVoteButton stays for the panel). And the card now shows the crowd split before the viewer answers, reversing the earlier call to withhold it. That is worth a second look on review: seeing "65% agree" first nudges the answer and makes the tally partly a measure of itself. Drawn as designed because it is a deliberate choice about what the card is for — a running read of the room rather than a poll.
Verified in a real browser against the live testnet debate, not just in jsdom: one card at 0:20, two stacked at 0:24 (the second a factual claim, correctly reading "100% verify" with chevrons rather than thumbs), and the paused state with both top-left controls up. 484 files / 5683 tests green, typecheck clean apart from the four files already failing on master.
One thing to eyeball rather than review in the diff: the card is capped at the 209px the frame draws it at. That is exact on the 484px explore card; on the much wider fullscreen player it stays 209px rather than scaling, on the theory that a card which grows to hold a paragraph stops being a glance. Easy to switch to proportional if it reads too small there.
Duplication, in the order it mattered:
Four graph ids were written out by hand in both publish scripts, and two of
them — `Types` and `Debate videos` — already existed in `core/debates/ontology.ts`
under different names. The other two, `Selector` and `Target property`, are
ontology ids with nowhere to live, so they live there now as
`SELECTOR_TYPE_ID` and `TARGET_PROPERTY_ID`. Both scripts import all four. The
emitted plan's constants are byte-identical, checked against a real run.
`arg()`, the four-line `--flag value` reader, had three copies. One, in
`scripts/lib/debate-claims.ts` beside the other shared script plumbing.
Dead surface: `DebateTicker`, `MAX_STACKED_CARDS`, `CHAT` and `toDashedUuid`
were exported and imported by nothing. Un-exported rather than deleted — each
is still used inside its own module.
`docs/plans/2026-09-15-…-matcher.py` is gone. It was the prototype that became
`core/debates/claim-timing.ts`, the only Python in the repo, and referenced by
nothing including its own sibling doc.
`scripts/dump-match-windows.ts` was in the same unreferenced state, but it is
the tool that produced the 59-claim read behind the 0.40 confidence bar, so
`claim-timing.ts` now cites it where it cites `calibrate-match-score.ts`. The
number's evidence is reproducible either way.
Coverage: `use-claim-timings.ts` had none, despite four rules worth holding —
including that a *failed* transcript counts as ready. Without that, a surface
ordering by time waits forever on a debate whose transcript will not load.
Also an unused `index` in the ticker's map, and a comment in
`use-debate-claim-response.ts` that explained why the claim page does not reuse
it but not why the explore card does not. Both leave `offersDebate` at its
default, which is the one thing that hook fixes in place.
Verified: production build passes, full suite 498 files / 5999 tests, eslint
clean across the files this PR touches.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…O-2909)
Four artefacts, 2,709 lines: two sets of agent instructions, a backfill plan
and the generated JSON it was executed from. They were working documents for a
one-off publish, not something the codebase needs to carry.
Three code comments referenced them and would have been left pointing at
nothing. Each keeps its reasoning and loses the link: the matcher's note about
where a better answer comes from now cites GEO-2958 alone, and the plan
script's header states the hand-done first debate as a fact rather than a
citation.
The one thing in those docs the code depended on was the shape of an answers
file, which only `build-plan-from-matches.ts` reads. That contract now lives in
its header, beside the parser it describes, which is where it should have been.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five of six findings were real. Each is fixed at the class, not the instance.
**A claim stated in two turns carried one turn's answers.** `find-or-create`
links an existing claim rather than minting a second, so one entity can be two
statements — the grouping tests call that the ordinary case. This PR then hung
`blockId`, `publishedTiming` and `relationEntityId` off the deduped row, where
the first turn read wins: a card over the wrong face, a timecode from the wrong
moment, and a backfill that writes one statement and abandons the other.
Measured first: 854 claims across all 82 debates, 854 block→claim statements,
zero restated. So the flaw is latent, and modelling timing per statement is
GEO-2958's job. What ships here is that it can no longer be silently wrong —
`restated` is set at the grouping, the resolver gives such a claim no timing at
all, and both backfill scripts skip and report it. It stays in the panel and
leaves the live layer, which is exactly where a claim whose speaker cannot be
resolved already goes.
**`Number(arg(…))` on a typo gave NaN, and every comparison against NaN is
false.** A mistyped `--floor` did not fail, it disabled the confidence floor and
planned a write for every match. Five sites had the shape — `--floor`,
`--limit` twice, and two positional arguments — so the fix is one validated
reader in the shared script module, used by all of them. A bad flag now exits
with a message.
**Duplicate answers were silently resolved to whichever came last**, because
`new Map(answers.map(…))` keeps the final entry for a repeated key — so the
"answered more than once" rejection the script documents could never fire. Now
detected while building the map. Swept the other nine `new Map(x.map(…))` sites
in the PR's reach: the rest key identical objects, where collapsing is correct.
**`verify-claim-timing` printed a negative count**: `assertable` spans every
source and a published timing scores 1, so subtracting it from the segment
count gave 0 − 13 on a fully published debate. Counted directly now.
**The claims chip was `hidden` on pointer devices, which took it out of the tab
order** — and since the corner draws nothing when no claim is live, a keyboard
had no way into the backlog at all. It is `absolute opacity-0` now: present,
focusable, costing no layout, and brought back into the flow on focus so the
ring lands on something visible. Hidden by position rather than `sr-only`
because that utility zeroes padding and height and Tailwind emits its
`not-sr-only` counterpart after `px-2` and `h-5` — checked against the compiler,
not assumed.
The sixth, bottom-left versus bottom-right, is the PR description being stale:
Preston asked for the mirror. One code comment still described the pre-mirror
arrangement and now does not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…areas (GEO-2909)
Two classes, each swept rather than patched at the reported site.
**The stopword list dropped one half of a polarity pair and kept the other.**
`not` was dropped while `never`, `no`, `cannot`, `nothing`, `neither` and
`without` were all scored, so "X is safe" and "X is not safe" reduced to the
same words and the matcher was free to place a claim over its own opposite.
Scanning the list for the same shape found two more: `more`/`most` dropped
against `less`/`fewer`/`least` kept, and `many` dropped against `few` kept.
All four now score, and the rule is written above the list so it does not drift
back. Measured over the 852 matchable claims: 54 windows move, 35 of them on
claims containing a negator, and the aggregate barely shifts — 131 scores rise,
147 fall, 574 unchanged, 13 newly clear the 0.40 bar and 7 newly fall below it.
Cheap in score; the 54 moves are the point. Modals stay dropped, all eight of
them, because dropping a whole class breaks no pair.
**An invisible control that still takes clicks.** The claim markers set
`pointer-events-auto` on themselves — they must, so a drag can pass between them
to the range input — and an explicit value beats the inherited `none` from the
hidden scrubber's wrapper, so a fully transparent marker stayed clickable and a
click meant to pause the video seeked it. Gated at the wrapper with a descendant
selector, which outranks the marker's own class. Focusability is untouched:
tabbing to a marker is what reveals the scrubber.
The sweep found a worse instance of the same class that had not been reported.
`recede` fades the mute button out once the viewer turns the sound on. On a
pointer device the hover that reveals it is the same gesture that makes a click
possible, so there is no window — but a phone has no hover, so the button sat
invisible and still tappable: a tap meant to pause muted instead, and nothing
could bring the control back. It no longer recedes where there is no hover.
Copilot also flagged the backlog chip as clickable while invisible. It was not —
`pointer-events` is inherited and the host sets `none`; measured, clicks at the
chip's coordinates land on the card behind it. But that depended on an ancestor
nobody editing the file can see, so it now says `pointer-events-none` itself.
**`verify-claim-timing` hard-coded 270s** while accepting any debate id, so
markers were scaled against the wrong timeline and sampling stopped at 4:30. It
reads the length off the transcript now. The default debate is 270s, which is
why it never showed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All five were real. None had been reported before; Copilot's "previously
missed" means it found them late, in code that had not changed.
**The backlog was built from the live layer's filtered list.** `tickerWindows`
keeps only claims firm enough to assert, and `historyBySlot` derived from the
same array — so the chip counted "N claims" and the list opened on a subset,
dropping everything the matcher had placed only loosely. `backlogWindows` is
its own list now: the live layer asserts "said at this moment" and needs a firm
one, the backlog only says "already said", which a whole-turn fallback answers
perfectly well once the turn ends. Invisible today — the backfill has since
published offsets for 853 of 854 claims — and true of every debate recorded
until GEO-2958.
**Disabled tickers fanned out row queries.** `useDebateTranscriptClaims` is
gated on `enabled`, but the feed card and the explore card fetch the same query
ungated, so a disabled ticker read a full claim list out of the warm cache and
handed it to `useDebateClaimsBySpaces`, which has no gate — authenticated row
queries and a gateway subscription per mounted feed item rather than for the
active pair. Empty groups while disabled.
**The answer map latched on the first side pressed.** Switch from Agree to
Disagree, or clear it, and the map kept the original — the scorecard would tally
an answer nobody holds. Every change reports now, `null` included, and the map
drops the entry. Returning the same map when nothing changed is what lets every
card report on mount without a re-render apiece.
**A marker seeked to where its card is invisible.** The hash sits at the claim's
end, which is the first frame of the card's fade-in, where opacity is zero — and
the scrubber is mostly used while paused, so the playhead stayed there and the
click showed nothing. Markers carry a `seekMs` just inside the window; `atMs`
still places the hash.
**And the markers could not be clicked at all.** The transparent range input
covers the whole bar at `z-2` against the marker layer's `z-1`, so every marker
click went to the input. Verified by hit-testing both stacks: the marker was
unreachable, and at `z-3` it takes its own click while the rest of the bar still
scrubs. The comment above it claimed the opposite, which is how it survived.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ale comments (GEO-2909)
Four findings, all real.
**The final marker seeked past the end of the recording.** `seekMs` lands 250ms
inside the card's window so the click shows a drawn card — but the player takes
the whole corner down once playback finishes, so on a debate whose last claim
ends at the final frame the offset produced exactly the blank it was added to
prevent. Measured across the 51 debates with published offsets: the median has
45 seconds of tail after its last claim, and **three end on it**. Clamped to a
millisecond short of the timeline.
Those three keep a residue no clamp fixes: a claim finishing as the recording
does has no frame its card could be drawn on, because the window opens where the
timeline stops. The seek lands *before* it rather than after, which draws no card
either way but leaves the corner up, so the chip and backlog are still there.
**Transcript failures were indistinguishable from missing transcripts.** A 5xx,
a dropped connection or an unparseable body all became an empty transcript —
whereupon the planner counts every claim of that debate as unmatched and writes
a plan that looks complete, the export drops the debate from its task files, and
the verifier prints "no transcript" and exits 0. An outage mid-run produced a
confident, quietly partial answer. Only a 404 means "not served" now; everything
else throws, and no caller swallows it, so a run stops before writing.
**Two comments had drifted from the code**, which is its own class and the one
worth sweeping. The chip's function doc still described the keyboard dead end as
intentional, contradicting the `className` three lines below that fixes it —
duplicate documentation, so the duplication went rather than being re-synced.
And three comments asserted a *live corpus state* ("853 of 854 published",
"zero of the 854 are restated", "140 of the 153"), two of which already
contradicted each other because they were measured weeks apart. Those now state
the rule and the direction the number moves; the counts that remain elsewhere
record a past measurement, which is a different thing and does not rot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…bility gate, loud script failures (GEO-2909)
Three findings, all real.
The marker seek clamp added last round was wrong. It held the seek a
millisecond short of the duration, but the player calls the last 50ms the
end and takes the whole claim corner down there — so on the three of 51
debates whose last claim ends on the final frame it produced exactly the
blank it was added to prevent. The threshold now lives in playback-utils
as PLAYBACK_END_EPSILON_MS, shared by playbackEnded, the background
recovery and the markers, and the test asserts against it rather than
against the duration.
Markers were built from every timed claim, while two later gates drop
some of them: a claim whose speaker is not in the participant list, and
one with no space to answer in. Both left a clickable hash and a chip
count with no card behind them. Rather than repeating the condition at
the marker site, the hook now filters once into renderableClaims and
feeds the live cards, the backlog and the markers from it, so the three
cannot disagree. Nothing in the corpus trips this today — measured, all
854 published claims have a space and an authored block — so this is for
the debates recorded from here on.
Same sweep as last round's transcript fix, applied to every remaining
place a failure could read as absence: export-claims-for-matching
swallowed a failed claims request as "no claims in this space" and still
exited 0 with a task set that looked complete, plan-claim-timecodes
counted load failures but exited 0 on an incomplete plan, and
build-plan-from-matches reported an unreadable answers file as a missing
one, quietly discarding the reader's work.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…backlog (GEO-2909)
Clicking the expand toggle on a live claim card threw the backlog open and
re-laid the claim out at the list's narrower width, so asking for the rest
of a sentence moved it into a different column mid-read. Pressing a thumb
on a live card did the same.
The stack reported every focus to the player, which opens the backlog on
it so a keyboard can reach the cards — but a click focuses whatever
control it lands on, and that read as having tabbed in. It now reports
only keyboard focus, using `:focus-visible`, which is the browser's own
answer to "did this focus come from a pointer" and the signal the backlog
chip already uses to decide when it is reachable.
Measured in Chrome inside the focus event rather than assumed: a click on
either control reports false, a Tab onto one reports true. jsdom models
both the same way, so the three tests hold the behaviour.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…script inputs (GEO-2909)
A marker for a claim ending on the final frame was still a broken
promise, and the comment defending it was wrong. The seek has to stay out
of the stretch playbackEnded calls the end, which is earlier than such a
claim's window opens — and both the live stack and the backlog are
bounded by `playheadMs >= startMs`, so the claim is in neither, wherever
the click lands. I had written that the clamp "at least leaves the corner
up so the backlog is still there to read it in"; spokenSoFar says
otherwise. Those claims now get no hash at all, which is the same answer
this function already gives a match it cannot place confidently.
Measured: 3 of 51 debates lose exactly one hash each, all three of them
hashes that showed nothing.
Four more places where a bad input produced a plausible-looking result:
- fetchTranscriptSegments mapped any 200 without a segments array to "no
transcript", so an error payload or a schema change read as "never
recorded" — the last door open in the sweep two rounds ago.
- arg() treated a flag with no value as an absent flag, so a trailing
`--limit` quietly meant every debate and `--out --limit 5` took
"--limit" for the output directory.
- export-claims-for-matching left an earlier run's task files in the
output directory, and build-plan reads every JSON in there — so a
rerun with a lower limit planned writes from stale data. It now clears
only files it recognises as its own output, by shape, so an answers
file sharing the directory is left alone.
- build-plan's duplicate-turn guard ran against writes as they
accumulated, which was too late: the first occurrence was already
pushed, or had been counted as confirmed/declined and pushed nothing at
all, so the claim could be reported rejected and planned anyway. It is
a pre-scan now, like the answered-twice guard beside it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#2449 rewrote the player's control cluster, which this PR had also been
editing.
Resolution:
- Took #2449's cluster wholesale. Both sides had independently added a
persistent play/pause beside mute; theirs is the designed one, with the
mobile centred glyph, the click-feedback flash and the larger controls.
- Kept `no-hover:opacity-100` on their mute rule alongside their
`md:opacity-100`. The two ask different questions: a tablet in
landscape is wider than `md` and still has no hover, so width alone
leaves the control faded but tappable there — the bug this PR fixed.
- Kept this PR's claim props on both tiles, the seam subtitle, the
extracted `useOpenDebaterProfile`, and the shared playback-end epsilon;
kept master's URL-fetch key fix and `releaseVideo` effect.
- The auto-merge dropped master's `showReplay`/`showPausedGlyph` and
re-introduced a `votes` prop and `WinnerVoteButton` that this PR's tile
redesign had removed. Both caught by diffing the merged file against
master's and accounting for every line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The four "previously missed" findings — all four real, all four fixed in 7002199f3
All of the same family, and the family is the one this PR has been sweeping for three rounds: a bad input or a failure producing a result that looks fine.
Malformed transcript responses treated as empty — fixed
Right, and this was the door left open by the fix two rounds ago. I made the status path honest and left the body path exactly as it was, so ?.segments ?? [] still turned any 200 that wasn't a transcript — an error payload, a schema change — into "this debate was never recorded". The contract said 404-only; the code had two ways out.
The body is validated now: anything without an array-valued segments throws with the debate id, and only a 404 still returns [].
A flag present with no value — fixed
Confirmed, both shapes you name. A trailing --limit quietly meant "every debate", and --out --limit 5 took "--limit" for the output directory and would have created one. Same shape as the NaN floor numberArg was written for, and on the same scripts — the guard was one level too high.
arg now refuses a missing value and refuses a next token that starts with --, exiting non-zero:
$ bun scripts/plan-claim-timecodes.ts --limit
--limit needs a value (exit 1)
$ bun scripts/export-claims-for-matching.ts --limit --out /tmp/x
--limit needs a value; got "--out" (exit 1)
Stale task files across reruns — fixed
Real, and worse in one direction than you describe: the plan is not just built from stale task data, it is built from what the graph looked like on the earlier run, so a claim published in between is re-planned from the old state.
The export clears them now, but only the ones it recognises as its own — each candidate is read, parsed, and required to carry a debateEntityId and turns. An --out pointed at a directory holding the reader's answers must not eat them, and I checked that it doesn't: planting an answers-shaped JSON and an unparseable file in the output directory, a rerun reports cleared 1 task files from an earlier run and leaves both standing.
Duplicate claims surviving the guard — fixed
Confirmed, including the second half, which is the worse one: a first occurrence that resolved to confirmed or declined pushes no write at all, so writes.some(...) never fired and the duplicate was never even reported.
It is a pre-scan now, built exactly like the answeredTwice set it sits beside — both occurrences dropped, one rejection line each rather than one per occurrence. Neither is trustworthy: two turns disagree about where the claim was said and nothing here can tell which is right.
Worth saying that the export should not be able to produce this — a claim carries one blockId — and measured over the 66 task files on disk there are none. That is the argument for making the guard correct rather than removing it: it exists for a case nobody expects, so it has to work the one time it is needed.
Verification: typecheck at master's baseline (5 pre-existing files, none in this PR), lint clean, 114 files / 2192 tests across the debate surfaces, and all four scripts smoke-run end to end. Also rebased onto master and merged #2449, whose player-control rewrite overlapped this PR's.
…umeric input (GEO-2909)
The polarity rule, through a door the stopword fix did not cover.
`wasn't` tokenised as one opaque word — neither `was` nor `not` — so a
claim's `not` had nothing to match, and the negated window carried a
spare content word that cut its precision. The affirmative therefore
outscored the negation:
claim "vaccination was not safe"
vs "vaccination was safe" 0.6667
vs "vaccination wasn't safe" 0.6111
which is the matcher free to place a claim over the speaker saying the
reverse. The corpus makes it one-directional rather than a coin toss: no
claim contracts (the extractor writes "does not"), while 248 of 7,633
segments do, and 90 claims say "not" in a turn that contracts it.
`n't` now expands to a `not` token before scoring, with curly apostrophes
folded first — `wasn’t` does not survive tokenisation at all, splitting
into `wasn` and `t`. Measured across the 852-claim corpus: 49 scores
rose, 0 fell, 23 windows moved, 6 claims newly clear the 0.40 floor and
none newly fall below it. Block-to-segment matching is unaffected, both
sides going through the same tokeniser.
Blank numeric input, swept across all three coercions this PR owns:
- `numberArg`/`numberAt`: `Number('')` is 0, which every range admits, so
`--floor "$FLOOR"` with FLOOR unset disabled the confidence floor
rather than failing — the exact outcome numberArg exists to prevent.
`arg` rejects a blank value too, so `--out ""` cannot name a directory.
- `publishedTiming`: a blank `integer` became offset 0, and 0 is a legal
start. A blank start beside a real end published "said in the first two
seconds" as a certainty, scored 1.0, over a claim made anywhere in the
debate. Blank is now absent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The comment promises a 28px touch target, but md: is a width query in this project, not an input query. A no-hover tablet wider than 767px therefore gets only a 20×20 button, despite the new no-hover variant specifically covering this hybrid case. Apply the larger target under no-hover as well.
…nd marker hit areas (GEO-2909)
Two more doors into the polarity fault, both in the same tokeniser.
`cannot` has no apostrophe for the n't rule to find, and it ran in both
directions: a claim saying "cannot" scored 0.667 against "can be safe"
and 0.611 against "can't be safe", while a claim saying "can't" scored
0.667 against "can be safe" and 0.611 against "cannot be safe". Whichever
spelling each side happened to use, the affirmative won.
STOPWORDS holds `it`, `that`, `we`, `there`, `i` but not `it's`,
`that's`, `we're`, `there's`, `i'm` — the apostrophe makes them different
tokens, so they survived as content words, and only on the speech side
since the extractor does not contract: `it's` 281 times in the corpus,
`that's` 143, `i'm` 82. Each was a distinct window word no claim could
match, cutting that window's precision — the same one-directional penalty
on speech that sounds like speech. Possessives fold the same way, and
want to: `person's` is `person`.
Measured across the 852-claim corpus: 218 scores rose, 3 fell, 44 windows
moved, 4 newly clear the 0.40 floor and none newly fall below. All three
that dipped kept the same window, so nothing was re-placed.
Scrubber markers were 2x10px targets, reachable on touch once a tap
pauses and reveals the bar. Both bounds on widening them were measured
rather than guessed: 34% of adjacent marker pairs in the corpus sit
closer than 24px on a phone, so a flat 24px target would cover the next
claim a third of the time; and these sit above the range input, so
whatever they cover is a place a drag cannot start — 18.5% of the track
at 24px, 10.4% at 12px. Each target is now 24px tall, up to 12px wide,
and clamped to the gap to its nearest neighbour, so two can touch but
never overlap. Hit-tested in Chrome: each marker wins its own centre,
the input still wins 20px away, overlap is zero. Focus is now visible
too; it was not before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When enabled is false, React Query can still return this debate's cached claims, so this computation still emits markers while groupBySlot deliberately returns empty live/history maps. An inactive feed card can therefore expose a clickable scrubber hash that seeks but can never show the promised card or backlog. Gate markers on enabled, just like the other ticker outputs.
The backlog chip is a sibling of this focus boundary. When a keyboard user opens the backlog from a live card and tabs from the list to the chip, relatedTarget is outside currentTarget, so this reports false; with a live card present the parent then closes the backlog and removes the chip before it can be activated, dropping keyboard focus. Include the chip in the same focus-within boundary (for example, handle focus/blur on the outer stack wrapper).
The reason will be displayed to describe this comment to others. Learn more.
Right, and chasing it found the bigger fault it was the visible edge of. My 2px floor only made the button pressable at all; it never asked why two buttons were at the same point.
They are at the same point because two claim entities carry the same text and the same published moment. So the marker overlap is one of four symptoms, not the problem:
two identical cards in the backlog, the same sentence twice
the chip counting both
two markers at one spot, the later covering the earlier
and an answer on one leaving the other unanswered
Measured over the corpus, and it is not hypothetical:
duplicate-text claims sharing a moment: 11 (all in one debate)
duplicate-text claims at different moments: 0
All eleven are in "Scot Loyd vs. CptMoh on Getting married…", whose claims appear to have been published twice — distinct ids, distinct blocks, identical text, identical offsets:
236000ms id=1e4f645b… block=5572c210… "Individuals are in a better position to benefit…"
236000ms id=f0646be0… block=d5e2b7c1… "Individuals are in a better position to benefit…"
So rather than grouping at the marker layer, the collapse goes in renderableClaims — the one gate the cards, the backlog, the count and the markers all read from, so none of them can learn about this separately and disagree again.
The key is text and moment, not text alone, and the zero in that table is why: a debater who repeats themselves later has said something new and should keep their card. Only the pair a viewer cannot tell apart on screen is collapsed.
One thing the collapse cannot fix, worth stating rather than hiding: the twin is a separate entity with its own responses, so an answer lands on whichever copy survived here. That is an argument for de-duplicating on publish, not for showing the same sentence twice — and the panel still lists both.
The reason will be displayed to describe this comment to others. Learn more.
Following up on my own reply above, because the de-duplication I described does not close this finding, and I only found that by checking rather than assuming.
Collapsing identical claims handles the 11 duplicate-text claims in the corpus. But those were measured from published offsets. The live matcher places a claim on segment boundaries, so any two claims matched to the same window end on the same millisecond — and those are different claims, correctly kept, still drawing two hashes at one point.
Measured over the corpus as the matcher places it:
assertable matched claims: 563
sharing an endMs with another claim: 16
of those, differently worded (still overlapping): 8
Not exotic at all — for example, both of these end at 96,400ms in one debate:
- The real aim of voter ID laws is to suppress voting in an attempt…
- There is significant evidence that voter ID laws suppress voting
So claimMarkers now emits one marker per moment, carrying a count, and the label says so rather than announcing one claim while silently standing for the others:
aria-label = count > 1 ? `Jump to ${count} claims, starting with: ${text}` : `Jump to: ${text}`
One hash per moment is the truer model regardless. The stack shows one card at a time, so seeking there surfaces the newest of them and leaves the rest a scroll away in the backlog — which is what the hash was promising in the first place.
Verified end to end over the 66 task files rather than only in unit tests:
markers drawn: 540
claims folded into a shared hash: 15
hashes still sharing a point: 0
The 2px floor in markerHitWidth stays as a guard, but nothing in the corpus reaches it now.
The reason will be displayed to describe this comment to others. Learn more.
This was fixed in 42516a8, shortly after the review that raised it — claimMarkers now collapses claims that finish on the same millisecond into one hash, so two targets on one point can no longer be produced. Reading the test as a permission was fair, though, so both ends are clearer now:
claim-ticker.test.ts holds the invariant at the producer — claimMarkers never emits two markers with the same fraction, and claimMarkers is the only thing that builds a ClaimMarker[], so holding it there holds it everywhere.
The markerHitWidth floor test now says it is a floor on an input the producer no longer generates, kept because the function is exported and pure and a 2px button that can be pressed beats a 0px one that cannot.
The grouped hash also announces its group (Jump to N claims, showing: …) and, as of 6d6a6c5, names the claim the card will actually show rather than the first of the group.
… chip focus (GEO-2909)
The coincident-marker report turned out to be the visible edge of a
bigger one. Two claim entities can carry the same text *and* the same
published moment, and then everything downstream doubles: two identical
cards in the backlog, a chip counting both, and two markers at one spot
where the later covers the earlier so the first cannot be reached with a
pointer. Measured: 11 such claims, all in one corpus debate whose claims
appear to have been published twice, and zero cases of the same text at a
different moment — which is why the key is text *and* moment rather than
text alone. A debater who genuinely repeats themselves later has said
something new and keeps their card.
Collapsed in `renderableClaims`, the single gate, so the cards, the
backlog, the count and the markers all agree rather than each learning
about it separately. What it cannot fix is that the twin is a separate
entity with its own responses, so an answer lands on whichever copy
survived — an argument for de-duplicating on publish.
Markers were not gated on `enabled`. `useDebateTranscriptClaims` is, but
the feed card and the explore card fetch the same query ungated, so a
disabled ticker reads a warm cache; `groupBySlot` refused to build cards
from it and the markers did not, leaving an inactive feed card able to
draw a hash that seeks to a claim whose card and backlog are both
deliberately empty. Gated in the same one place.
The stack's focus boundary sat on the scroll box, which is the chip's
sibling, so tabbing from the last card out to the chip reported focus
gone: the player closed the backlog, `showChip` went false with a live
card present, and the chip unmounted mid-tab, dropping focus to the body.
The one control that reopens the list was the one unreachable from inside
it. The boundary is the whole stack now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two "previously missed" findings — both real, both fixed in c36a6b517
Markers not gated on enabled — fixed
Correct, and it is the half I left behind when I fixed the gating in the third round. useDebateTranscriptClaims is gated, but the feed card and the explore card fetch the same query ungated, so a disabled ticker reads a full claim list out of a warm cache. groupBySlot refused to build cards from it; the markers never asked. An inactive feed card could therefore draw a clickable hash that seeks to a claim whose card and backlog are both deliberately empty — the same broken promise this PR has been closing everywhere else.
Gated in renderableClaims rather than at the marker site, so it is one gate covering the cards, the backlog, the count and the markers together. groupBySlot keeps its own enabled check, now redundant, with a comment saying so — it takes windows as an argument and a future caller could hand it some from elsewhere.
Backlog chip outside the focus boundary — fixed
Confirmed, and the consequence is worse than a stale flag: it drops focus entirely.
The handlers sat on the scroll box, which is the chip's sibling. Tabbing from the last card out to the chip put relatedTarget outside currentTarget, so the boundary reported focus gone; the player cleared focusedSlot, showChip went false because a live card was present, and the chip unmounted underneath the keyboard mid-tab. Focus fell to the body, and the one control that can reopen the list was the one that could not be reached from inside it.
The boundary is the whole stack now, chip included. The :focus-visible guard moved with it unchanged, so a click still does not read as tabbing in.
There is a test, and I confirmed it bites by narrowing the boundary back to the scroll box — it fails with expected last "vi.fn()" call to have been called with [ true ], which is exactly the moment focus was reported gone.
Verification: typecheck at master's baseline (5 pre-existing files, none in this PR), lint clean, 114 files / 2210 tests across the debate surfaces. All four tests added this round were confirmed to fail without their fix.
The de-duplication in the previous commit handled identical claims, but
not the case that turns out to be ordinary: `matchClaimWindow` places a
claim on segment boundaries, so any two claims matched to the same window
end on the same millisecond. Measured over the corpus as the matcher
places it rather than from published offsets, 8 of 563 assertable claims
share a moment with a differently-worded claim — two hashes at one point,
the later covering the earlier, the first unreachable by pointer.
`claimMarkers` now emits one marker per moment, carrying a `count`, and
the label says so rather than announcing one claim while silently
standing for others. One hash per moment is the truer model anyway: the
stack shows one card at a time, so seeking there surfaces the newest and
leaves the rest a scroll away in the backlog.
Verified end to end over the 66 task files: 15 claims fold into a shared
hash, and no two hashes share a point.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five commits, overlapping this PR's player changes again.
- Took master's `awaitingTap` (userPaused || autoplayBlocked) rewrite of
`showControls` and `showPausedGlyph`, and its `data-debate-*`
diagnostic attributes on the player root.
- Kept this PR's ticker hook, and its single-surface root className: no
gap and one radius, which is what lets the subtitle straddle the seam
rather than sit inside one tile.
Checked the same way as the last merge — diffed the merged file against
master's and accounted for every line — and confirmed each of master's
additions landed: awaitingTap, the effect's dependency change, and all
four data attributes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
Agreed, and this was the one worth the most here — it fails silently and is invisible afterwards. Every check in this loop is against the task file as it is now, so a re-cut transcript passes all of them and resolves the same indices to different milliseconds. The script keeps its promise that a timecode is a real boundary of the recording and publishes the wrong boundary.
Task files now carry a taskVersion — a 12-hex fingerprint of what an answer actually depends on: the turn a claim is filed under, the index-to-milliseconds mapping, and the claim text. Deliberately notpublished or matcherGuess, which are re-read from the current task file when the plan is built, so a backfill landing mid-read does not throw away correct work. It lives in scripts/lib/debate-claims.ts so the exporter and this script cannot drift on what a version is.
The answers echo it, and a mismatch rejects the whole file rather than individual claims — the drift is a property of the transcript, so every answer read against the old one is equally suspect. They fall back to the matcher, which is where they would have been with no answers file at all. A task file predating the stamp fails the same way.
Verified end to end against fixtures rather than by reading:
matching pair → placed by reading: 1 (0–2400ms)
re-cut, stale answers → rejected: 1
! Getting married: answers were read against c52681999fc1, this task file is 5e2f743448f0 — re-read it
unversioned answers → rejected: 1
That middle case is the bug: the old code accepted it and planned 0–2600ms. taskVersion has its own suite in scripts/lib/debate-claims.test.ts covering what it does and does not notice.
I checked the rest of the PR for the same shape — a handoff between two processes validated by name alone. This is the only one: plan-claim-timecodes.ts builds its plan from live data in one pass, and live-claim-reuse.ts takes a single explicitly-named payload.
The reason will be displayed to describe this comment to others. Learn more.
Right, and reproduced before fixing. MAX_STACKED_CARDS is 1, so .slice(-max) is .slice(-1) and the card that comes up is the last of the group while the hash was labelled with the first.
A test that seeks to the marker and asserts the card matches it fails on the old behaviour:
AssertionError: expected 'b' to be 'a'
The collapse now keeps the last of the group as the representative, so id/text are the claim the click surfaces. Both sides sort stably on the same key over the same array (windowsFor and claimMarkers both sort by the claim's end, and the reachability filter is constant within a group), so "last" is the same element in both. The aria-label for a group changed from "starting with" to "showing" to match, and the grouping test now asserts b.
… (GEO-2909)
Three findings, all real.
A grouped scrubber hash named the wrong claim. `claimMarkers` collapses claims
that finish on the same millisecond into one hash and kept the *first* of them
for its label and tooltip; `tickerStack` keeps one card and takes it from the
end of the group, so clicking the hash announced one sentence and put a
different one on screen. The hash now carries the claim the card will show, and
a test seeks to the marker and asserts the card that comes up is the one the
label named — it fails on the old behaviour.
Answers read against an earlier cut of a transcript could publish the wrong
offsets. `build-plan-from-matches.ts` pairs answers to tasks by filename and
claim id, and everything it checks — the claim exists, the indices are in the
turn, the range runs forwards — still passes after a re-export whose segment
boundaries moved. The script's whole promise is that a timecode is a real
boundary of the recording, and that failure keeps the promise while publishing
the wrong boundary. Task files now carry a `taskVersion` fingerprint of what the
reader read; the answers echo it, and a mismatch throws the file out to the
matcher rather than planning from it. Verified end to end against fixtures: a
matching pair plans 0–2400ms, the same answers against a re-cut transcript
(which would have published 0–2600ms) are rejected, and an unversioned answers
file is rejected too.
The third was already fixed by the hash collapse in 42516a8 — two targets on
one point can no longer be produced. The invariant now has a test at the
producer, and the hit-width floor says it is a floor rather than a permission.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
Read through this at the PR head. The timing resolver's three-source design is sound, and the four findings in the description all check out against the code — the scoreWindow precision fix and the scrubber-fraction correction especially.
Four inline notes below, one of them worth fixing before merge. Plus one question that isn't a line-level defect:
Winner voting is gone from the tile, and its replacement is in #2459.WinnerVoteButton and the position_label chip are removed from DebaterVideo, but the player still branches on hasVoted for showControls and showReplay. Standalone, that leaves an end-of-debate state offering a replay to a viewer whose only remaining way to vote is inside the claims panel. If these merge in order it's a non-issue; if this one can land alone, worth deciding that deliberately rather than by merge order.
The reason will be displayed to describe this comment to others. Learn more.
focusedSlot can latch open when the stack unmounts with focus inside it.
The only thing that clears it is onFocusChange(false) from the stack's onBlur, and that handler is the sole caller — there's no unmount cleanup. But claimsFor returns null (unmounting the stack) when playbackEnded, when the ticker is disabled because the card scrolled out of the active/preload window, or when cards and history are both empty.
Browsers don't fire blur/focusout for a focused element removed from the DOM, so a viewer who tabs into the stack and then reaches the end of the debate leaves focusedSlot set for the life of the player. On replay, claimsOpenFor is true immediately and that debater's corner opens straight into backlog mode with no pointer or keyboard in it — and nothing closes it short of tabbing back in and out.
pinnedSlot has the same shape on touch, where the only clear is the mouse-filtered pointerleave.
The reason will be displayed to describe this comment to others. Learn more.
Confirmed and fixed — thank you, this was the serious one.
Reproduced both halves before touching it. The ticker suite now asserts your premise directly: focus inside the stack, unmount, and onFocusChange is never called with false. The player suite asserts the consequence, and both of its tests fail without the fix.
I went with clearing at the owner rather than a cleanup inside the stack, for two reasons: the useEffect(() => () => onFocusChange?.(false), []) form captures a stale prop under an empty dep array, and it leaves pinnedSlot — which has the identical hole on touch, as you noted. So claimsFor's mount condition is lifted into stackShownFor(slot) and an effect releases both latches for any slot whose stack has gone:
That covers all three unmount paths you listed — playbackEnded, the ticker emptying when the tile leaves the active/preload window, and cards and history both empty. Tests cover the first two.
The reason will be displayed to describe this comment to others. Learn more.
filled is passed to icons that don't accept it, so a factual claim never gets the filled selected state.
ChevronUp / ChevronDown take { color? } only and drop the prop; ThumbUp / ThumbDown take { color?, filled? } and use it to fill the glyph.
So on a veracity claim the selected Verify/Dispute chevron is distinguished only by bg-white/15 plus text-green / text-red-01, where a stance claim also fills. Minor, but it's a real inconsistency in the one affordance that says "you have taken a side". Either give the chevrons a filled variant or stop passing it.
The reason will be displayed to describe this comment to others. Learn more.
Confirmed — ChevronUp/ChevronDown take { color? } and drop it; ThumbUp/ThumbDown take { color?, filled? }. The union-typed Icon const is what let it through the compiler.
I have taken the half that belongs in this PR and split the half that does not.
Here: stop passing it. The shared Icon const is gone and each glyph is now called with the props it actually has, so a future filled on a chevron has to be a deliberate edit rather than a silent no-op:
Not here: giving the chevrons a filled variant. ~/design-system/icons/chevron-up.tsx is shared across the app, and Preston has asked that fixes outside the feature a PR touches get reported rather than made in it. So this one is a report: the veracity selected state is bg-white/15 + text-green/text-red-01, where stance gets that and a filled glyph. I am not sure it is a bug — a chevron is a stroke with no interior, so "filled chevron" may not be a real design — but the inconsistency you describe is there, and it is in the one affordance that says the reader has taken a side. Worth a design call, and a separate PR if the answer is yes.
The reason will be displayed to describe this comment to others. Learn more.
This points at an override that isn't there — the card applies [mask-image:var(--claim-ramp,none)] unconditionally, and the only no-hover: in the file is on the chip. Looks like it survived the touch exception being removed (which the card's own doc block explains). Worth deleting the second sentence; in a file this dependent on its comments it'll send the next reader hunting for a guard that doesn't exist.
The reason will be displayed to describe this comment to others. Learn more.
You are right — the override is gone and the comment outlived it. The card sets [mask-image:var(--claim-ramp,none)] unconditionally at line 362, and the only no-hover: left in the file is on the chip.
Rather than just delete the sentence I replaced the whole rationale, because the first half leaned on the same missing guard ("so the card's own stylesheet can refuse it"). There is still a real reason for the custom property, and it is the one visible in the className:
// Through a custom property rather than `mask-image` itself: the card declares the standard
// and `-webkit-` masks together in its own className, so one variable drives both and the
// declarations stay with the rest of the card's styling.
The reason will be displayed to describe this comment to others. Learn more.
blocks isn't deduped by id, although the claim loop below it is.
The doc block's reasoning for deduping claims — "the graph can return the same relation twice (duplicate publishes do happen)" — applies just as well here, but blocks.push runs once per block relation with no seen check.
The app is unaffected: resolveClaimTimings and the ticker both key blocks into a Map. The scripts don't. export-claims-for-matching.ts iterates claims.blocks directly and filters claims.all by blockId, so a duplicated block relation emits the same turn twice with the same claims — which then trips the claimedTwice guard in build-plan-from-matches.ts and rejects every claim in that turn as "appears in more than one turn of this task file".
Worth noting that guard's own comment says "The export should not produce this — a claim carries one blockId". This is the path that makes it produce exactly that, and the failure is silent: dropped placements, not an error. Deduping on push is two lines.
The reason will be displayed to describe this comment to others. Learn more.
Confirmed, including the failure path — good catch, and the silence is the worst part of it.
Verified the signatures: blocks.push runs once per block relation with no seen check, while the claim loop below dedupes through rowsByClaimId. A test that feeds the same block relation twice fails on the old code with two identical entries in blocks.
Worth noting the claim loop handles the repeat correctly on its own — row.blockId !== blockEntity.id is false for the same block, so restated stays false, which is what its comment says it should do. The hole was only the turn list.
Deduped on push, with the reasoning pointed at the consumer rather than the symptom, since nothing in the app would ever show it:
The test asserts both halves — one turn listed, and the claim still not marked restated.
On the guard comment in build-plan-from-matches.ts you quoted: I have left it saying the export should not produce this, because with this fixed that is true again, and the guard's own reason for existing is that it has to be correct rather than plausible.
…rop (GEO-2909)
Four findings from @ohohoreilly. Three were real bugs; the fourth was a comment
pointing at a guard that no longer exists.
The backlog could latch open for the life of the player. `focusedSlot` is
cleared only by the stack's own `onBlur`, and a focused element removed from the
document fires no blur — so tabbing into the backlog and then reaching the end
of the debate, or scrolling the tile out of the preload window, left the slot
latched. On replay the corner opened straight into backlog mode with nothing in
it, and nothing short of tabbing back in and out closed it. `pinnedSlot` had the
same hole on touch, where the only other release is a mouse-filtered
`pointerleave`. Both are now released when the stack goes. The ticker's suite
asserts the premise directly — an unmounted stack reports nothing — and the
player's asserts the fix, for the end of the debate and for the ticker emptying.
Turns were not deduped, though the claims beside them were, and the doc block's
reasoning ("the graph can return the same relation twice") applies equally. The
app never saw it because every reader keys blocks into a Map; the matching
scripts walk the list, so a repeated block emitted one turn twice with the same
claims in each, and the plan builder read that as a claim filed under two turns
and dropped every claim in it — a silent loss of placements from a duplicate
this layer exists to absorb.
`filled` was handed to icons that drop it: the chevrons take `color` alone, so a
veracity claim never got the fill a stance claim does. The prop is no longer
passed to them. Whether veracity's selected state wants more than its background
and colour is a design-system question, raised separately rather than answered
here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GEO-2909 · spec: Claims in Motion · backend follow-up: GEO-2958
Claims are a debate's payload and they lived behind a button, in a side panel, ordered by ranking score, with no relation to what was on screen. This puts them where the attention is.
What you get
Claims appear over the video once they have been said. Each debater has their own corner in the bottom-right of their own tile, built to the Figma frame: a glass card carrying their avatar and name, the crowd's share, the claim clamped to two lines, and two small icons — thumbs for a stance claim, chevrons for a factual one.
One corner per debater rather than one for the player, because a viewer is looking at whoever is talking; a shared corner asks the eye to leave the speaker in order to read what the speaker is saying.
A card appears at the claim's end, not its start. Anchored to the start it arrived while the debater was still mid-sentence — asserting a claim the viewer had not heard them make yet, and asking them to agree with something still being argued.
The corner rests on a chip, and opens into the backlog. Cards expire, so the corner is empty for most of a debate; a small "N claims" pill says the backlog is there. Pointing at that debater's tile, tabbing into the stack, or pressing the chip opens everything they have said so far — scrollable, newest at the bottom, dissolving into the tile's edge. Bounded by the playhead, so it is a record rather than a table of contents.
The chip exists because hover is nothing on a touch screen and close to nothing on a desktop — an invisible affordance is only found by accident. On touch the chip is deliberately the only way in: the tile's hover handler ignores anything but a real mouse, or the synthesised
pointerenterfrom a tap would throw the corner open every time someone tapped to pause.Markers on the scrubber, one per precisely-placed claim, each a seek — at the claim's end, where its card appears.
The claims panel is in spoken order. It sorted by ranking score, which is right for a feed of unrelated claims and wrong inside a transcript: a debate is an argument, and reading it out of sequence loses the thread. Ties — every claim of one turn — still fall to the ranking. Rows whose moment is firm enough carry a timecode.
How it knows when a claim was said
core/debates/claim-timing.ts— one resolver, three sources in order:publishedsegmentblockSo a debate published long before any of this still gets an answer, and a better one the day offsets land for it. Both reads are shared cache entries, not new requests: claims come back on the key the count badge and panel already use, the transcript on the key the player already fetches for subtitles.
One bar decides whether a moment may be stated.
isAssertableMoment— a timecode and a card over the video both go through it. They used to disagree: the live layer asked for 0.55 while the panel's timecode asked only that the claim had been matched at all, so a 0.40 match judged too loose to draw over the video still printed a time to the second, indistinguishable from a published offset. Ordering deliberately does not go through it: a list has to put a claim somewhere, and where it lands asserts far less than a timestamp.Four findings worth a look
A block's markdown is the verbatim concatenation of its Whisper segments, so locating a turn is string equality, not fuzzy matching. All six blocks matched exactly. That is why the
blockfallback is trustworthy rather than a guess.Relation
positionis not transcript order.debate-publish-draft.tsassignsPosition.generate(), which is random rather than monotonic. On the test debate the blocks in position order are the 3rd, 1st, 6th, 4th, 5th and 2nd turns — so the orderingtranscript-claims.tsdocumented as "transcript order" was arbitrary on every debate since it shipped. Fixed, and the panel now orders by recovered time.The scrubber markers were misplaced. Their fractions were scaled by the latest claim's end while the track is sized by the debate's timeline. Those agree only when the last claim runs to the final second — true of the hand-published test debate and false elsewhere. Across eight live debates the gap reached 11%: about 20 seconds out on a 210-second debate, on a control whose whole job is landing you somewhere exact.
scoreWindowcould exceed 1. A repeated claim word was counted once per occurrence against a set of window words, so its precision term could pass 1 and push the score past the [0, 1] it documents. Two claims scored 1.05 and 1.03; a claim saying "AI" three times outranked a tighter match saying it once.Verified against live data
bun scripts/verify-claim-timing.ts <debateEntityId> <spaceId>runs the real grouping, resolver and ticker selection over testnet. On the test debate all 13 claims resolvepublished; on one with none, the three tiers come apart as they should:Behaviour was checked in a real browser rather than only in jsdom — which is how the
line-clampoverride, the staleResizeObserver, and the top-anchored open list were all caught.Backfill
scripts/plan-claim-timecodes.tsproduced the plan the publishing agent ran: 140 writes across 50 debates at a 0.7 confidence floor. The floor is the point —publishedis the resolver's most-trusted source, scored 1.0, so writing a weak match launders a guess into an exact answer no later gate can demote.This has since been run: the graph now holds 153 published timecodes (13 hand-published + 140). Re-running the planner produces an empty plan, which is the idempotency working.
The planning documents and the generated plan JSON were working papers for that one-off publish and are not committed — Preston asked for them out of the PR. The scripts are self-describing and stay; the record of which offsets were written, and at what confidence, lives with the run rather than in the repo. Worth knowing for GEO-2958, which will want to know which offsets it may overwrite with the extractor's real spans.
Test data
Timecodes for the test debate were hand-published first so the
publishedpath had something real to read — thirteen claims, four of them corrected by hand after reading the transcript.Reusing the existing
Start offset/End offsetInteger properties rather than minting new ones was Preston's call and a good one — they belong to theSelectortype the graph already uses to anchor aReply toat a span of its target. Note the API serialises Integer values as strings; parsed and range-checked with a test per failure mode, because a half-parsed pair would be drawn as a real moment.Testing
claim-timing.test.ts— a verbatim turn of the real debate; each of its three claims must land on the sentence that says it. Real data, because the thing under test is whether an extractor's paraphrase matches back to speech, and a hand-written fixture would quietly be written to match. Plus the scorer's range, the assertable bar, and spoken-order sorting.claim-ticker.test.ts— the end-anchored window, stack capping, expiry, the backlog, fades, and marker placement against the timeline.debate-claim-ticker.test.tsx— attribution on the card, the vocabulary verb on the share, the expand toggle, the chip, and no play/pause leak.debate-claims-panel.test.tsx— spoken order beating the ranking, and the ranking still breaking ties.transcript-claims.test.ts— offset parsing, six discard cases.apps/websuite: 483 files, 5708 tests, green.tsc --noEmitclean apart from the four files already failing onmaster.Self-review pass
e7d1f87removes the duplication this PR introduced: three copies of the claims query collapsed to one shared by the app and both scripts, a hand-rolled traversal in the planner replaced by the relation entity id the grouping now carries, and theuseClaimResponseState+useClaimPositionControlwiring — repeated at five call sites, two of them added here — extracted intouseDebateClaimResponse. Net −224 lines, no behaviour change.Not in this PR
Nothing publishes offsets automatically yet. GEO-2958 covers geo-chat emitting
start_ms/end_msper extracted claim, with the rule that where the extractor knows the span it should report it rather than leave it to be guessed.Running it locally
The debates feed needs geo-chat, which
apps/web/.envdoes not point anywhere by default: