Skip to content

Latest commit

 

History

History
58 lines (31 loc) · 9.02 KB

File metadata and controls

58 lines (31 loc) · 9.02 KB

CreatorFlow — thinking document

The problem I chose

The brief describes a process: creators get a comped visit, we get raw footage. Building the process is the obvious read, and it is not where the pain is. Spreadsheets track visits badly, but they do track them. What nothing tracks is what came back.

The failure that costs real money and makes no noise is this. Footage arrives as IMG_4471.MOV, RPReplay_Final1721654433.mp4, trim.9F3A1C22.MOV in a shared drive. Somebody skims it, drags a few clips into an edit, moves on. Six months later an editor needs a vertical sauna shot with nobody in frame, cleared for paid use. They cannot find it — not because it does not exist, but because nothing about the pile is searchable. So they re-shoot footage the company already owns, or they use a clip they are not licensed for and nobody notices until legal does.

Nothing in that story looks like a bug. Every individual visit went fine. The loss is invisible because it is spread across time — which is exactly why a CRM is the right place to fix it. The moment to catalogue footage is the moment it arrives, while somebody still cares.

The second half only becomes visible once the first is solved: a searchable library still hands you files, and an editor is almost never looking for a file. They are looking for the second the water hits the stones.

The solution

The brief is the contract. Deliverables are structured rows, not prose — 2× establishing, cold plunge, vertical, minimum 6s. That one artifact does three jobs: it is what gets negotiated, what submitted footage is counted against, and how the library is organised. One thing to maintain instead of three that drift apart.

Ingest creates the value. When footage lands, a four-tier pipeline names every clip, describes it, tags it against a controlled vocabulary, grades it, checks it against the brief, and indexes what happens inside it second by second. The demo submission resolves to 5 of 9 delivered, and the four failures are each a different kind: a pool-deck clip standing in for the cold plunge; a file named like a talking head that is an empty corridor; a clip both soft and full of unreleased guests; and one that matches the brief but runs 4.1 seconds against a 6-second minimum. That last is the important one — the model matched it, the code disqualified it.

The library pays it back. Clips shelve by spa and by creator in either nesting, search in plain English, and search inside: "the moment the water hits the stones" returns 0:04.3 in a specific clip, and clicking it opens the player there. Then the library drafts the cut — which clips, in what order, which seconds of each — and where the shelf cannot cover the goal, the brief for the next visit. Which closes the loop back into the pipeline.

Filters decide eligibility; relevance decides order. Conflating those is how library search usually goes wrong. Taxonomy filters exclude, and appear as chips the editor can remove. The rest of the sentence only ranks — against the descriptions and per-frame labels written at ingest, plus the spa and the creator, because "Nadia's San Jose plunge stuff" is how people ask. Ranking never excludes: matches on top, everything else below a divider, each card naming the words it matched. The failure being designed against is the empty grid, which an editor cannot tell apart from "we do not own this".

Where I used AI, and where I deliberately did not

The principle: AI proposes, code disposes, a human approves.

AI does what is genuinely judgement-shaped and expensive for a person: reading a frame and saying what is in it, matching a clip to a requirement, translating a sentence into filters, scoring a creator against an anchored rubric, drafting a message, reading a coverage matrix, choosing what should follow what in a cut.

Code does everything with a right answer. All arithmetic — package values, deliverable counts, duration minimums, runtime totals. The state machine: a model never mutates CRM state. The attention queue: date rules with exact reason strings. Rights: human-entered structured fields, code-computed expiry, snapshotted onto each asset at publish, so amending an agreement cannot retroactively re-license footage already in a campaign. A hallucinated "cleared for paid" is a legal problem, not a UX one.

The useful version of this is the split inside one feature. The vision pass writes a label per frame; the timestamp comes from the frame extraction, never the model. The edit planner chooses the beats; the runtime is summed in code, out-points are clamped to real durations, and anything that overran is visibly marked as trimmed. Rights and orientation are filtered before the planner runs, so an ineligible clip is not discouraged from the plan — it is never offered.

Two cuts on purpose. Embeddings: a vector match that surfaces the wrong clip gives an editor nothing to correct; a chip saying "Scene: Pool deck" costs one click. Worth revisiting past ~10k assets. An AI morning digest: the signals are date arithmetic, and wrapping them in a model adds latency and hallucination surface to information that has to be exactly right.

One constraint shaped everything: the Claude API does not accept video. Vision means images, so "analyse the footage" necessarily means "analyse extracted frames" — keyframe selection is not a shortcut, it is the whole game. Frames are chosen on scene change rather than fixed intervals, the prompt tells the model it is looking at stills and to hedge on anything motion- or audio-dependent, and the UI shows which frames it saw so you can check it.

Decisions worth defending

Token economics as a design constraint, not an optimisation. Four tiers, each filtering for the next: free ffmpeg measurement, cheap triage, vision on three frames, text-only reasoning. Nine clips cost ~30k tokens; the naive one-call-with-thirty-frames version costs about fifteen times that for a worse answer. The moment index rides along free — the frames are already in context, so one extra line per frame is ~60 output tokens. A second pass would have cost the whole vision tier again.

The taxonomy is a closed enum in every AI schema. The model physically cannot return a tag the database does not know. Taxonomy drift — the thing that quietly kills every media library — is prevented structurally rather than by asking nicely.

Demo mode is a first-class path. With no API key every feature still works, and the deterministic paths are computed from the real data rather than replayed, so they stay correct when a reviewer approves a clip or picks a filter nobody anticipated. It doubles as the failure path when a live call errors.

The demo data has real defects, not labelled ones. Where the scenario says a creator delivered one soft clip, the build genuinely softens that file and the pre-filter measures a real defect. Labelling a sharp clip "blurry" in a fixture would make the whole quality pipeline decorative. That principle cost the most time: the softness metric took three attempts, because edge energy measures brightness and contrast-normalised edge energy measures scene detail — a person against smooth water scored below a deliberately blurred clip. What works is comparing each clip against a blurred copy of itself, which cancels brightness, contrast and subject alike.

How I prioritised

Seed-data first: the whole demo narrative worked off seeded rows before a single live AI call, so the product thinking was testable on day one and the AI was built against a UI that already existed. Then the two "wow" paths — submission review and library search — then connective tissue, then hardening.

Cut consciously: real file storage, auth, tests, e-signature, a creator dashboard. The one overspend was real video, and I would make it again, because a prototype that asks you to imagine the footage is asking you to imagine the product. It grew a second time for the same reason: a twenty-clip library can be browsed, so search never has to work. Building the back-catalogue out to where scrolling stops being an option is what turned search from a feature into the point.

What I would build next

  1. Video-native analysis for the claims stills cannot settle — continuous takes, audio, "does the shot hold". Gemini takes video natively. Keep tiers 0 and 1 and reserve the expensive pass for clips where the brief hinges on something a frame cannot show. It is also what makes the moment index dense rather than three points per clip; transcribing talking heads puts spoken words on the same timeline.
  2. Learning from editors. Every download and every "used in project" is a quality signal. After a few hundred, calibrate the grading rubric against what editors actually reach for rather than what the model finds pretty.
  3. Batch backfill of the existing archive — the biggest single win available, at half price through the Batch API.
  4. Repeat-partner mechanics. Nadia is on her third visit. Repeat creators are cheaper and better than new ones, and nothing currently optimises for that.