A small, self-trained recommendation model for Nuvio — the actual model
half of what nuvio-track's README correctly flagged as unbuildable from
one person's data alone: "This can't be bootstrapped from one person's
data... not true collaborative filtering."
This is the fix for that: a small two-tower retrieval model, pre-trained
on public collaborative data (MovieLens 25M for movies, TMDB "similar
shows" weak supervision for TV) so it's useful to a brand new user from
first install, not just after you personally have watch history.
Integration into Nuvio itself lives in a separate branch of the app repo
(feature/on-device-recommender on
NuvioMobile) — this repo is
just the model: data pipelines, training, export, and eval.
See NuvioMedia/NuvioMobile#1803 for the feature-request writeup this was built for.
This is the standard architecture behind most real-world recommenders (YouTube's recommender, TensorFlow Recommenders, etc.), just at a much smaller scale here:
- Item tower: embeds a movie/show — genre multi-hot, normalized
release year, a movie/TV type flag, and hashed word embeddings of its
title + TMDB overview + keywords (see
combine_text()inmodel/two_tower.py) — into a d=64 vector. Word hashing, not a learned per-item lookup table: any title, in or out of the training set, hashes into the same fixed-size table, so the item tower can embed a catalog item it's never seen before without retraining. - User tower: mean-pools the item-tower embeddings of a user's recent watch history into the same d=64 space.
- Training objective: in-batch sampled softmax (InfoNCE-style) — for each (user, item-they-liked) pair, the model is trained so the user vector is closer to that item's vector than to every other item in the same training batch. This needs no negative-sampling infrastructure and is the standard way to train retrieval two-towers.
At inference, recommending for a user is just: encode their history once,
then dot-product against all candidate item vectors and take the top-K.
Runs in well under 50ms on CPU — no GPU needed for serving, only for
training. The actual on-device integration goes further: the item side
never runs on-device at all, since it's precomputed once for the whole
catalog offline (export_weights.py) — only the small user tower needs
to run live.
- Movies: MovieLens 25M (25M ratings, 62K movies) for real
collaborative "people who liked X also liked Y" signal, enriched with
TMDB overview + keyword text (
data/enrich_tmdb.py) for real content signal beyond just genre/year tags. - TV: there's no MovieLens-scale public dataset of real people's TV
ratings, so TV shows get weak supervision instead — TMDB's own "similar
shows" relationships (
data/fetch_tv_pairs.py) become item-item positive pairs, trained with the same contrastive objective as the movie side but comparing two items directly instead of a user against an item.
Nearest-neighbor sanity check, cosine similarity in the trained item embedding space — comparing title-only (v1) against the final model (title + overview + keywords, plus the TV item-item pass):
| Probe | v1 (title-only) | Final |
|---|---|---|
| The Matrix | The Sender, Nemesis 3, Replicant — obscure same-genre/year clones | Fight Club, Saving Private Ryan, The Sixth Sense |
| Silence of the Lambs | Henry: Portrait of a Serial Killer, After Midnight | Forrest Gump, Braveheart, Seven |
| The Shining | Same-decade horror obscurities | Psycho, Jaws, A Clockwork Orange, Apocalypse Now |
Title-only mostly learned to cluster on genre+year tags — technically
correct neighbors, but obscure ones. Real overview/keyword text clearly
picks up thematic content instead, surfacing recognizable, plausible
recommendations. Run it yourself: python3 eval.py --checkpoint model/checkpoints/two_tower_v2_latest.pt.
A real limitation found the same way: a handful of TV shows show up
as false-positive "matches" for movies they have nothing to do with (the
item-item TV pass has no explicit signal keeping TV embeddings distant
from unrelated movie embeddings — see the comment on recommendFor() in
the app-side NuvioRecommender.kt for the serving-side stopgap currently
in place, and why it isn't a real fix). Worth fixing in training before
this goes further than a demo — an explicit cross-type negative term, or
scaling back the TV pass relative to the movie pass, are the two obvious
next things to try.
source .venv/bin/activate
cp .env.example .env # fill in a TMDB API key: https://www.themoviedb.org/settings/api
python3 data/prepare_movielens.py # download + ingest ml-25m
python3 data/enrich_tmdb.py # movie overview/keywords (rate-limited, ~1.3hr for all 62K)
python3 data/fetch_tv_pairs.py # TV catalog + similarity pairs (~20min for 8K shows)
python3 train.py --epochs 5 # trains model/checkpoints/two_tower_v2_latest.pt
python3 eval.py --checkpoint model/checkpoints/two_tower_v2_latest.pt
python3 export_weights.py --checkpoint model/checkpoints/two_tower_v2_latest.ptTrained on an Apple M2 Mac Mini (16GB RAM, PyTorch MPS backend) — full 5-epoch run takes roughly 1.5-2 hours end to end including data prep. Deliberately not fedora, which runs a live Minecraft server and was already fighting a real OOM problem the same night this was built.
- MovieLens ingest, TMDB movie enrichment, TV catalog + similarity pairs — all memory-safe streaming, no pandas DataFrame or Python list-of-lists held in RAM for the big tables
- Two-tower model, content-based item tower (title/genre/year/type + overview/keywords), no closed vocabulary
- Training: movie user-item objective + TV item-item objective, shared item tower, 5 epochs
- Export to the plain-binary format the on-device Kotlin inference
code loads (
export_weights.py) - Qualitative eval (
eval.py) — see Results above - Fix the cross-type leakage properly (in training, not just the serving-side penalty currently in place)
- Personalization pass on real
nuvio-trackwatch data (blocked on that actually having real usage) - A real recall@K number against held-out interactions, not just nearest-neighbor spot checks