Skip to content

docs(pitch): honest F10 narration — recalibration, not a fake flip#90

Merged
ComBba merged 1 commit into
mainfrom
pitch/honest-f10-narration
Jun 6, 2026
Merged

docs(pitch): honest F10 narration — recalibration, not a fake flip#90
ComBba merged 1 commit into
mainfrom
pitch/honest-f10-narration

Conversation

@ComBba

@ComBba ComBba commented Jun 6, 2026

Copy link
Copy Markdown
Contributor

The live /judge cohort recording shows 'no rank change in this cohort', AUDIT EFFECT ±0 pts, hit@13 Δ=0 — so the old F10 line 'the audit changes who wins' would directly contradict the on-screen footage (the exact overclaim the 3-agent review flagged as highest-risk).

Rewrote F10 to match the real footage: the audit recalibrates every score against the evidence, the ranking holds, Δ=0, 'the audit doesn't fake a flip — it earns the score.' This turns the honest Δ=0 into a credibility point (stronger than a flashy fake flip).

  • Per-beat F10 + continuous take updated.
  • Added a code comment guarding against re-introducing 'changes who wins' over that footage.

Pairs with the cleaned /judge clip (loading 15s→1s) for the video's F10 segment.

Summary by CodeRabbit

릴리스 노트

  • 문서
    • 나레이션 스크립트의 주요 대사 부분을 업데이트하여 증거 기반 점수 재보정 프로세스 및 순위 변화 없음을 더 명확히 전달하도록 수정했습니다.
    • 화면 영상과의 일치성을 확보하기 위해 관련 설명을 개선했습니다.

The live /judge cohort footage shows 'no rank change in this cohort', AUDIT EFFECT
+/-0 pts, hit@13 delta=0 — so the old F10 line 'the audit changes who wins' would
contradict the on-screen result (the exact overclaim the review flagged). Rewrote F10
to match the real footage: the audit recalibrates every score against the evidence,
the ranking holds, delta=0, 'the audit doesn't fake a flip — it earns the score'.
Turns the honest Delta=0 into a credibility point. Per-beat + continuous take updated
+ a code comment guarding against re-introducing the overclaim.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Jun 6, 2026

Copy link
Copy Markdown

Wondering what really moved? Review this PR in Change Stack to inspect semantic changes, definitions, and references.

Review Change Stack

개요

pitch/narration-script.md의 F10 "the payoff" 내레이션 구간을 업데이트했습니다. 기존의 "감사가 누가 이기는지 바꾼다"는 표현을 제거하고, 증거 기반 재보정과 순위 유지(델타 0)를 강조하는 내용으로 교체하여 화면 영상과의 일관성을 확보했습니다.

변경사항

F10 내레이션 일관성 업데이트

레이어 / 파일 요약
F10 'The Payoff' 스크립트 개정
pitch/narration-script.md
F10 구간의 내레이션을 세그먼트별 대사(56-63줄)와 연속 테이크(73줄)에서 동시에 갱신했습니다. 기존의 "누가 이기는지 바꾼다"는 문장을 제거하고, Arize에서 모든 span(6 hats/48 evidence spans)을 확인하며 점수를 증거 기반으로 재보정하고 델타 0(순위 변화 없음)을 보여준다는 내용으로 교체했습니다. 화면의 "no rank change" 시각화와 모순되지 않도록 하는 HTML 정직성 주석을 추가했습니다.

관련 PR

  • Two-Weeks-Team/glasshat#81: 두 PR 모두 pitch/narration-script.md의 피치 내레이션을 수정합니다(관련 PR은 F1–F11 전역의 스크립트를 갱신했고, 이 PR은 F10 "the payoff" 구간을 특별히 수정하여 순위 변화 표현 제거).

예상 검토 난이도

🎯 1 (간단함) | ⏱️ ~3분

축하 시

🐰 토끼가 소곤거려 봅니다—
"내레이션과 화면이 춤을 추네요,
증거 재보정으로 정직한 이야기,
델타 제로, 순위는 뉘앙스로,
피치의 진실이 반짝반짝!" ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed 제목은 PR의 주요 변경 사항을 명확하게 요약합니다. F10 대사를 '거짓된 플립'에서 '재보정'으로 변경하는 것이 핵심이며, 제목이 이를 직접적으로 반영합니다.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pitch/honest-f10-narration

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the narration script in pitch/narration-script.md for the F10 segment to reflect an honest recalibration rather than claiming the audit changes who wins, ensuring the narration matches the actual video footage. The review feedback correctly identifies that these updates need to be synchronized with the corresponding slide script in pitch/deck.html to avoid discrepancies, and suggests adding missing pause markers to the full continuous take to maintain consistent pacing.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread pitch/narration-script.md
`F10` · the payoff · honest recalibration · speed 1.0
```
Open Arize. Every span is there. <#0.35#> And on a whole cohort, the audit changes who wins.
Open Arize. <#0.3#> Every span is there — six hats, forty-eight evidence spans under one judgment. <#0.4#> Nothing hidden. <#0.6#> Now run the whole cohort. <#0.3#> Every score, recalibrated against the evidence. <#0.5#> Here, the ranking holds — and we show it: delta, zero. <#0.4#> The audit doesn't fake a flip. <#0.3#> It earns the score.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The updated narration script for F10 is not synchronized with the corresponding slide script in pitch/deck.html (Slide 11 / F10). Specifically, pitch/deck.html lines 1134–1142 still contain the old heading (<h2>The audit changes who wins.</h2>), the old voiceover text ("Open Arize. Every span is there. And on a whole cohort — <b>the audit changes who wins.</b>"), and outdated stage delivery notes. Please update pitch/deck.html to match this new honest recalibration narration to prevent discrepancies during the presentation.

Comment thread pitch/narration-script.md
## Full continuous take — paste as ONE (≈1,050 chars, under the 5,000 limit) · speed 0.97
```
AI writes the submissions now. By the thousands. Every hackathon, every grant, drowning in them. <#0.35#> But the judging didn't change. Still a score nobody audits. <#0.6#> An AI judge is worse. Confident. Fast. Wrong. <#0.45#> With no way to check it. <#0.6#> So we flipped it. Don't only judge the work. <#0.3#> Audit the judgment itself. <#0.8#> Trace it. <#0.6#> Trust it. <#1.0#> How? <#0.3#> Every hat. Every retrieval. Every score. <#0.35#> A span in Arize AX. A hundred and four of them. One evaluation. <#0.6#> One ADK 2.0 Workflow. Six hats in parallel. Deployed as a real agent on Agent Engine. <#0.4#> Not a notebook. <#0.6#> It reads the evaluator's own rules. Synthesizes the rubric. <#0.3#> Then every hat grounds its score in retrieved evidence. <#0.6#> Six perspectives, in parallel. <#0.3#> Watch Yellow. Optimism. <#0.3#> Run hot. <#0.6#> Now watch. Yellow over-scored. <#0.3#> Against a held-out calibration prior, strongest where the evidence is thin, <#0.4#> the audit catches its own over-confidence. And pulls the score back. <#0.55#> Live. With the math shown. <#0.6#> Open Arize. Every span is there. <#0.35#> And on a whole cohort, the audit changes who wins. <#0.8#> Open source. Apache 2.0. <#0.5#> Trace it. <#0.45#> Trust it.
AI writes the submissions now. By the thousands. Every hackathon, every grant, drowning in them. <#0.35#> But the judging didn't change. Still a score nobody audits. <#0.6#> An AI judge is worse. Confident. Fast. Wrong. <#0.45#> With no way to check it. <#0.6#> So we flipped it. Don't only judge the work. <#0.3#> Audit the judgment itself. <#0.8#> Trace it. <#0.6#> Trust it. <#1.0#> How? <#0.3#> Every hat. Every retrieval. Every score. <#0.35#> A span in Arize AX. A hundred and four of them. One evaluation. <#0.6#> One ADK 2.0 Workflow. Six hats in parallel. Deployed as a real agent on Agent Engine. <#0.4#> Not a notebook. <#0.6#> It reads the evaluator's own rules. Synthesizes the rubric. <#0.3#> Then every hat grounds its score in retrieved evidence. <#0.6#> Six perspectives, in parallel. <#0.3#> Watch Yellow. Optimism. <#0.3#> Run hot. <#0.6#> Now watch. Yellow over-scored. <#0.3#> Against a held-out calibration prior, strongest where the evidence is thin, <#0.4#> the audit catches its own over-confidence. And pulls the score back. <#0.55#> Live. With the math shown. <#0.6#> Open Arize. Every span is there — six hats, forty-eight evidence spans under one judgment. Nothing hidden. <#0.5#> Now run the whole cohort. Every score, recalibrated against the evidence. <#0.4#> Here, the ranking holds — and we show it: delta, zero. The audit doesn't fake a flip. It earns the score. <#0.8#> Open source. Apache 2.0. <#0.5#> Trace it. <#0.45#> Trust it.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The F10 segment in the full continuous take is missing several pause markers (<#0.3#>, <#0.4#>, etc.) that are present in the per-beat F10 script (line 58). Aligning these pause markers ensures that the continuous voiceover generation maintains the same pacing and rhythm as the individual per-beat clips.

Suggested change
AI writes the submissions now. By the thousands. Every hackathon, every grant, drowning in them. <#0.35#> But the judging didn't change. Still a score nobody audits. <#0.6#> An AI judge is worse. Confident. Fast. Wrong. <#0.45#> With no way to check it. <#0.6#> So we flipped it. Don't only judge the work. <#0.3#> Audit the judgment itself. <#0.8#> Trace it. <#0.6#> Trust it. <#1.0#> How? <#0.3#> Every hat. Every retrieval. Every score. <#0.35#> A span in Arize AX. A hundred and four of them. One evaluation. <#0.6#> One ADK 2.0 Workflow. Six hats in parallel. Deployed as a real agent on Agent Engine. <#0.4#> Not a notebook. <#0.6#> It reads the evaluator's own rules. Synthesizes the rubric. <#0.3#> Then every hat grounds its score in retrieved evidence. <#0.6#> Six perspectives, in parallel. <#0.3#> Watch Yellow. Optimism. <#0.3#> Run hot. <#0.6#> Now watch. Yellow over-scored. <#0.3#> Against a held-out calibration prior, strongest where the evidence is thin, <#0.4#> the audit catches its own over-confidence. And pulls the score back. <#0.55#> Live. With the math shown. <#0.6#> Open Arize. Every span is there — six hats, forty-eight evidence spans under one judgment. Nothing hidden. <#0.5#> Now run the whole cohort. Every score, recalibrated against the evidence. <#0.4#> Here, the ranking holds — and we show it: delta, zero. The audit doesn't fake a flip. It earns the score. <#0.8#> Open source. Apache 2.0. <#0.5#> Trace it. <#0.45#> Trust it.
AI writes the submissions now. By the thousands. Every hackathon, every grant, drowning in them. <#0.35#> But the judging didn't change. Still a score nobody audits. <#0.6#> An AI judge is worse. Confident. Fast. Wrong. <#0.45#> With no way to check it. <#0.6#> So we flipped it. Don't only judge the work. <#0.3#> Audit the judgment itself. <#0.8#> Trace it. <#0.6#> Trust it. <#1.0#> How? <#0.3#> Every hat. Every retrieval. Every score. <#0.35#> A span in Arize AX. A hundred and four of them. One evaluation. <#0.6#> One ADK 2.0 Workflow. Six hats in parallel. Deployed as a real agent on Agent Engine. <#0.4#> Not a notebook. <#0.6#> It reads the evaluator's own rules. Synthesizes the rubric. <#0.3#> Then every hat grounds its score in retrieved evidence. <#0.6#> Six perspectives, in parallel. <#0.3#> Watch Yellow. Optimism. <#0.3#> Run hot. <#0.6#> Now watch. Yellow over-scored. <#0.3#> Against a held-out calibration prior, strongest where the evidence is thin, <#0.4#> the audit catches its own over-confidence. And pulls the score back. <#0.55#> Live. With the math shown. <#0.6#> Open Arize. <#0.3#> Every span is there — six hats, forty-eight evidence spans under one judgment. <#0.4#> Nothing hidden. <#0.6#> Now run the whole cohort. <#0.3#> Every score, recalibrated against the evidence. <#0.5#> Here, the ranking holds — and we show it: delta, zero. <#0.4#> The audit doesn't fake a flip. <#0.3#> It earns the score. <#0.8#> Open source. Apache 2.0. <#0.5#> Trace it. <#0.45#> Trust it.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
pitch/narration-script.md (1)

58-58: ⚡ Quick win

UI 배지 문구와 내레이션 핵심 표현을 더 직접 정렬하는 것을 권장합니다.

현재 문구도 의미상 맞지만, 화면의 고정 카피(no rank change in this cohort)와 내레이션의 핵심 문장을 더 가깝게 맞추면 이후 카피 변경/재녹음 시 불일치 리스크를 줄일 수 있습니다.

Also applies to: 73-73

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@pitch/narration-script.md` at line 58, Adjust the narration copy to match the
on-screen badge text verbatim so there's no drift between UI and voice: replace
or append the current phrasing around "the ranking holds — and we show it:
delta, zero." and/or "The audit doesn't fake a flip. It earns the score." with
the exact badge copy "no rank change in this cohort" (or include that exact
phrase immediately after the existing line) and apply the same change to the
parallel line later in the script ("no rank change in this cohort" also used at
the other noted instance).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@pitch/narration-script.md`:
- Around line 57-59: The fenced code block containing "Open Arize. Every span is
there — six hats, forty-eight evidence spans under one judgment..." is missing a
language tag and triggers MD040; update the opening fence from ``` to include a
language identifier (e.g., ```text or ```plain) so the block is recognized as
plain text and the markdownlint warning is resolved.

---

Nitpick comments:
In `@pitch/narration-script.md`:
- Line 58: Adjust the narration copy to match the on-screen badge text verbatim
so there's no drift between UI and voice: replace or append the current phrasing
around "the ranking holds — and we show it: delta, zero." and/or "The audit
doesn't fake a flip. It earns the score." with the exact badge copy "no rank
change in this cohort" (or include that exact phrase immediately after the
existing line) and apply the same change to the parallel line later in the
script ("no rank change in this cohort" also used at the other noted instance).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 45463a18-ad3b-41e5-86a6-8fdf2683d2ef

📥 Commits

Reviewing files that changed from the base of the PR and between 68d85e0 and aecb5ab.

📒 Files selected for processing (1)
  • pitch/narration-script.md

Comment thread pitch/narration-script.md
Comment on lines 57 to 59
```
Open Arize. Every span is there. <#0.35#> And on a whole cohort, the audit changes who wins.
Open Arize. <#0.3#> Every span is there — six hats, forty-eight evidence spans under one judgment. <#0.4#> Nothing hidden. <#0.6#> Now run the whole cohort. <#0.3#> Every score, recalibrated against the evidence. <#0.5#> Here, the ranking holds — and we show it: delta, zero. <#0.4#> The audit doesn't fake a flip. <#0.3#> It earns the score.
```

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

코드 펜스 언어를 명시해 markdownlint 경고를 해소하세요.

Line 57 코드 블록에 언어 식별자가 없어 MD040 경고가 발생합니다. ```text 같은 언어 태그를 붙여 문서 린트 일관성을 맞추는 게 좋습니다.

🧰 Tools
🪛 markdownlint-cli2 (0.22.1)

[warning] 57-57: Fenced code blocks should have a language specified

(MD040, fenced-code-language)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@pitch/narration-script.md` around lines 57 - 59, The fenced code block
containing "Open Arize. Every span is there — six hats, forty-eight evidence
spans under one judgment..." is missing a language tag and triggers MD040;
update the opening fence from ``` to include a language identifier (e.g.,
```text or ```plain) so the block is recognized as plain text and the
markdownlint warning is resolved.

Source: Linters/SAST tools

@ComBba
ComBba merged commit 223a2ae into main Jun 6, 2026
5 checks passed
@ComBba
ComBba deleted the pitch/honest-f10-narration branch June 6, 2026 10:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant