Skip to content

Roadmap: H3 multishot era — battery results, unlock stack, and the next build wave #85

Description

@taskmasterpeace

Tracking issue for everything the multishot validation opened up (2026-08-06). Results update as tests land.

Learning battery (plan: docs/plans/2026-08-06-h3-multishot-learning-battery.md)

Test Question Status Result
Captain 30s Does chaining work at all? ✅ done PASS — 30.3s continuous, identity through both seams
T1 rap, 9:16 Musical performance + vertical ✅ rendered 30.3s @ 480×832 with audio — awaiting Robert's grade
T2 artist photo anchor REAL artist identity through a chain ✅ rendered Likeness PASS at 25s (frame-checked). ⚠️ FINDING: opening frames stretched — the pack resizes the portrait start_image to 16:9 with stretch mode (_resize(..., \"disabled\")); recovers as generation takes over. Fix for our integration: center-crop/pad the start image to target aspect before feeding.
T3 coverage (CU→wide→tracking) Reframing tolerance 🔄 rendering
T4 hard location cuts Scene-change tolerance (decides MV architecture) 🔄 queued
T5 60s six-shot verse Long-chain drift + runtime 🔄 queued
T6 voice reference (ref2va + captain's voice) Voice conditioning lever 🔄 queued

The unlock stack (researched, sources in repo memory)

  • H3 character LoRAs: ostris AI-Toolkit added H3 T2V/I2V LoRA training 2026-08-03 (NVFP4, consumer GPUs) — already installed here; Fizgig trains from still photos. → artist identity without reference images.
  • H3 native sing-lipsync claims (feed a singing voice → lip-synced video; re-lipsync existing footage). Untested locally — follow-up test after T6.
  • LTX-2.3 IC-LoRA family: Motion-Track-Control (official), Cameraman (camera-move transfer from reference videos), Canny/Depth/Pose transfer. → real MV camera language on generated footage.
  • LTX as H3 upscaler (community workflows; upscaler model already on disk).

Build wave (from Robert's review, 2026-08-06)

  • H3-aware Prompt Enhancer redesign: click enhance → LLM analyzes the prompt → offers category directions (pick one) → generates 4 full enhanced prompts, all visible, click-to-use. Bake the H3 prompting guide into the system prompt. Edge-case tests: very long prompts (5000 chars), empty prompt, unicode, malformed LLM output fallbacks.
  • Gallery: click-through lightbox — prev/next arrows + keyboard navigation instead of X-and-reopen.
  • Gallery: bulk erase — multi-select + one delete.
  • Prompt persistence everywhere — every render carries its prompt (.meta.json sidecars); battery outputs backfilled by hand this time.
  • Characters in Playground discoverability — the machinery EXISTS (ReferencePicker + @name mentions resolve library characters + quick modes) but reads as absent. Surface it (Character button chips like Gen Space).
  • Multishot into DD (plan in memory): generate_multishot on the h3 client + Playground Multishot toggle + crop-not-stretch start images + eventually Music Video sections as single chains.

Done this wave (context)

Music Video rename · compose-while-rendering · queue work surface (reorder/edit/duplicate/chains) · seed fix + seeds visible · Gallery duration/resolution/seed · credential split · model registry + honest ETAs + clapperboard.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions