Cross-lab QC for what the builder cannot perceive. Fable 5 (Claude) builds the UI and can read a still frame. Google Gemini 3.1 Pro watches the video and hears the audio.
A cross-lab QA gate for AI-built output. One model builds; a multimodal model from a rival lab grades the real rendered result against a rubric you control and returns a strict JSON verdict. The build model applies the highest-impact fix, the loop repeats until it passes, and a human signs off last.
Built live while an AI agent built a studio landing-page hero: Claude (Fable 5) wrote the React + three.js, and Google's Gemini 3.1 Pro was the judge.
The model writing the code can already see a screenshot. What it cannot do is watch the animation play or hear the track. Fable 5 (Claude) is multimodal for still images, so it can check a static frame on its own. But it cannot watch video or hear audio. So the moment the output is motion (a scroll animation, a three.js scene, a rendered reel) or sound (a music bed, a voice track), the builder is blind and deaf to its own work, and that output ships un-judged.
A still screenshot is the weakest case, the one thing the builder didn't actually need help with. The real gap is everything that moves or plays. Pairing the builder with a rival multimodal model closes it: Gemini 3.1 Pro watches the video and hears the audio, scoring the real result against a rubric you control. Fable 5 reasons and builds; Gemini perceives and judges. The two labs' models are complementary, not redundant, and that is the whole point: nothing motion-driven or audio ships un-checked, and every fail comes back with the single highest-impact fix.
The loop:
- Build. Your build model writes or edits the UI, the motion, or the audio.
- Capture. Capture the real output: a screenshot for static UI, a short clip for motion, the rendered file for audio.
- Judge. Feed the capture + a rubric to a rival-lab model (Gemini 3.1 Pro, via the Google
Antigravity CLI). It returns a strict JSON verdict
{pass, score, axes, top_fix, notes}. PASS only if the total clears the bar and no single axis fails. - Fix. The build model applies
top_fix, re-capture, re-judge. Repeat to PASS. - Human signs off last.
The judge runs through the Google Antigravity CLI on Gemini 3.1 Pro (High), which self-authenticates: no API keys are read or handled by this tool. (Antigravity also sidesteps the public API quota wall on the preview model.)
The included qc-hero.sh judges a screenshot of a rendered UI, the static base case:
# capture a screenshot of your running UI (any headless browser works), then judge it:
./qc-hero.sh shot.pngSwap RUBRIC-hero.md for your own definition of "world-class." The same gate is what you point at the
modalities the builder cannot perceive on its own: a video reel or an audio track, watched and
heard by a model that can actually see motion and listen. A verdict looks like:
{
"pass": true,
"score": 89,
"axes": {
"hierarchy": 9, "typography": 7, "whitespace": 9, "brand_coherence": 10,
"color_contrast": 9, "component_polish": 8, "motion_readiness": 9,
"composition": 10, "readability": 9, "agency_grade": 9
},
"top_fix": "Tighten tracking on the main display headline for a more premium, editorial feel.",
"notes": [
"Flawless strict left-edge alignment throughout the layout.",
"The teal particle atmosphere looks rich and sophisticated, not cheap.",
"Consider slightly darkening the inactive pill borders for clearer definition."
]
}