Video is evidence. Point it at raw video and this app builds a live Neo4j knowledge graph of what is shown, said, and written — then lets an OpenAI-brained agent answer questions over it, with a live graph visualization.
Uses all four: TwelveLabs (Marengo + Pegasus video understanding) · OpenAI (agent brain + entity canonicalization) · AWS Strands (agent + tool orchestration) · Neo4j (context graph + vector index + NVL viz).
The thesis: the same entity seen across many independent videos MERGEs to one node — so the graph grows richer instead of duplicating.
video-context-graph/
├── backend/ FastAPI + Strands (OpenAI) agent, TwelveLabs ingestion
│ ├── app/twelvelabs_client.py Marengo embed/search + Pegasus analyze
│ ├── app/agent.py OpenAI agent + graph tools (SSE streaming)
│ ├── app/vector_client.py Neo4j vector search over segment embeddings
│ ├── app/routes.py FastAPI endpoints
│ └── scripts/ingest.py video -> Pegasus -> OpenAI -> Neo4j pipeline
├── frontend/ Next.js + Chakra UI + NVL (chat | graph | video inspector)
├── cypher/schema.cypher constraints + vector index
├── data/ontology.yaml Video/Segment/Entity/Topic ontology
└── .env credentials (not committed)
(:Video)-[:HAS_SEGMENT]->(:Segment {embedding}) // vector index on Segment.embedding
(:Segment)-[:NEXT]->(:Segment) // temporal order
(:Segment)-[:MENTIONS]->(:Entity) // MERGE'd across videos (key = normalized name)
(:Segment)-[:ABOUT]->(:Topic) // MERGE'd across videos (key = normalized name)
Entity and Topic nodes are keyed by their normalized name, so the same
person/place/object appearing in two different videos collapses to one node —
that shared node is the whole point.
- uv (Python 3.10–3.13) — backend
- Node.js 18+ / npm — frontend
- Neo4j 2025.x+ — Aura free (
neo4j+s://…) or local Docker (make docker-up). A vector index is required, which modern Neo4j provides out of the box. - API keys: OpenAI and TwelveLabs (a free tier works).
cp .env.example .env # fill in NEO4J_*, OPENAI_API_KEY, TWELVE_LABS_API_KEY
make install # backend (uv sync) + frontend (npm install)
make seed # ingest the vendored sample clip -> build the graph
make start # backend :8000 + frontend :3000- Frontend: http://localhost:3000 — chat, live NVL graph, video inspector
- Backend: http://localhost:8000 —
GET /health,/api/...
Open the frontend, ask "What videos do we have and what are they about?", and watch the graph light up.
All ingestion goes through backend/scripts/ingest.py. make seed is a thin
wrapper around it. There are three ways to add a video, and you can mix them.
The repo ships a ready-to-use clip at data/videos/bbb_1080p_30fps_normal_85sec.mp4
(an 85-second excerpt of Big Buck Bunny — see Sample video & attribution).
Running make seed with no arguments ingests every .mp4 in data/videos/,
so a fresh clone builds a populated graph out of the box:
make seedDrop your own .mp4 files into data/videos/ and they'll be picked up the same way.
TwelveLabs downloads the file server-side, so the URL must be directly and publicly fetchable.
# via make (uses the same index as everything else)
make seed VIDEOS="https://example.com/a.mp4 https://example.com/b.mp4"
# or call the script directly
cd backend && uv run python scripts/ingest.py https://example.com/a.mp4The file is uploaded into your TwelveLabs index. After indexing, the app stores the HLS URL TwelveLabs returns so the clip still plays back in the UI.
cd backend && uv run python scripts/ingest.py /path/to/clip.mp4If you already indexed a video in TwelveLabs, ingest it straight into the graph by id — this skips upload/indexing and goes right to analyze → embed → write.
cd backend && uv run python scripts/ingest.py \
--index-id=<TL_INDEX_ID> --video-id=<TL_VIDEO_ID>The
index_id/video_idmust belong to the account behindTWELVE_LABS_API_KEYin your.env. Ids from a different account return403 read_not_allowed.
- Index the video with TwelveLabs (Marengo + Pegasus) —
tasks.create+wait_for_done(skipped for method 3). - Analyze with Pegasus → a rich, time-coded description.
- Structure that prose with OpenAI Structured Outputs → schema-validated segments, each with a summary, on-screen text, transcript, canonicalized entities, and topics.
- Embed each segment with Marengo (512-dim) for the Neo4j vector index.
- Write to Neo4j:
Video/Segmentcreated,Entity/TopicMERGE'd across videos, a temporalNEXTchain, and the vector index ensured.
Indexing a short clip takes ~1–2 minutes; longer videos take proportionally more.
- Format/access: direct MP4 over http(s). YouTube/Drive/S3-signed share
links do not work — TwelveLabs must be able to
GETthe raw file. (The oldcommondatastorage.googleapis.com/gtv-videos-bucket/*samples now return 403;https://test-videos.co.uk/...clips are known-good.) - Resolution: use 360p or higher — very small frames are rejected.
- Duration: at least ~4 seconds.
- See TwelveLabs' limits for exact size/duration bounds.
Re-running ingestion for the same video replaces that video's old segments
(and their embeddings) and re-MERGEs shared entities/topics — no duplicate
Video/Segment nodes. So you can tweak the pipeline and re-seed freely.
Ingest two different clips that share a subject (e.g. both feature the same
person, place, or object). Their shared Entity/Topic nodes become a single
node connected to both videos — visible as a hub in the graph.
👉 HOWTO.md is a step-by-step walkthrough of adding a first video, then a second, and exactly how/why the entities and topics merge.
make reset # ⚠️ DETACH DELETE every node in the Neo4j database
make schema # re-apply constraints + indexes only (no data)make reset wipes the entire database, not just this domain — don't run it
against a Neo4j instance you share with other data.
| Variable | Default | Purpose |
|---|---|---|
NEO4J_URI |
neo4j://localhost:7687 |
Bolt URI (Aura: neo4j+s://…) |
NEO4J_USERNAME / NEO4J_PASSWORD |
neo4j / — |
Neo4j auth |
NEO4J_DATABASE |
neo4j |
Target database |
OPENAI_API_KEY |
— | Agent reasoning + structured video extraction |
OPENAI_MODEL |
gpt-5.6 |
Strands agent reasoning and graph tools |
OPENAI_EXTRACTION_MODEL |
gpt-5.6-terra |
Schema-validated video and entity extraction |
OPENAI_REASONING_EFFORT |
low |
Keeps agent and extraction responses fast; also accepts none |
TWELVE_LABS_API_KEY |
— | Read by the TwelveLabs SDK |
TL_INDEX_ID |
(empty) | Reuse a specific index; else created/found by name |
TL_INDEX_NAME |
video-context-graph |
Index name when creating |
MARENGO_MODEL |
marengo3.0 |
Index + search model |
PEGASUS_MODEL |
pegasus1.2 |
Analyze model (index-creatable) |
MARENGO_EMBED_MODEL |
marengo3.0 |
Segment embeddings (512-dim) |
SAMPLE_VIDEO_URLS |
a Big Buck Bunny clip | Default clip(s) for make seed |
DOMAIN_ID |
video-context-graph |
Tags all nodes; keep consistent across ingest + app |
BACKEND_PORT / FRONTEND_PORT |
8000 / 3000 |
Ports |
CORS_ORIGINS |
http://localhost:3000 |
Comma-separated allowed origins |
| Role | Value | Note |
|---|---|---|
| Index + search | marengo3.0 |
multimodal (visual + audio + transcription) |
| Analyze | pegasus1.2 |
the Pegasus version an index accepts at creation |
| Segment embeddings | marengo3.0 |
512-dim text embeddings → Neo4j vector index |
| Agent brain | gpt-5.6 |
via Strands OpenAIResponsesModel, with low reasoning |
| Video extraction | gpt-5.6-terra |
OpenAI Structured Outputs with a typed segment schema |
Note: TwelveLabs accepts
pegasus1.2when creating an index butpegasus1.5is analyze-only. The embedding dimension (512 formarengo3.0) is auto-detected at ingest time and the vector index is created to match.
| Target | What it does |
|---|---|
make install |
Install backend (uv) + frontend (npm) deps |
make seed [VIDEOS="..."] |
Ingest videos (env sample, or the given URLs/paths) |
make schema |
Apply Neo4j constraints + indexes only |
make start |
Run backend + frontend together |
make dev-backend / dev-frontend |
Run one side |
make reset |
|
make test-connection |
Verify Neo4j connectivity |
make docker-up / docker-down |
Local Neo4j via Docker |
make test |
Backend unit tests, frontend type-check, and end-to-end test discovery |
make test-e2e |
Run end-to-end tests against the running, seeded application |
make lint |
Linters |
| Method + path | Purpose |
|---|---|
GET /health |
Backend + Neo4j status |
POST /api/chat |
One-shot agent turn |
POST /api/chat/stream |
Streaming agent turn (SSE) |
GET /api/videos |
List ingested videos + segment counts |
GET /api/videos/{id}/segments |
A video's segments in order |
POST /api/search |
Live multimodal search via TwelveLabs |
GET /api/schema · GET /api/schema/visualization |
Graph schema |
POST /api/expand |
Neighbors of a node (graph drill-down) |
POST /api/cypher |
Run a read Cypher query |
search_video_moments— embed the query with Marengo, vector-search Segments (find-the-moment)explore_graph— traverse everything involving an entity/topic/video across all videostwelvelabs_search— live multimodal Marengo search straight from TwelveLabsrun_cypher/get_graph_schema— arbitrary read-only graph access
| Symptom | Cause / fix |
|---|---|
media_url_not_accessible on ingest |
The URL isn't directly fetchable by TwelveLabs. Use a raw MP4 link (not a share/streaming page). |
403 read_not_allowed with --video-id |
That video/index belongs to a different TwelveLabs account than TWELVE_LABS_API_KEY. Use the matching key, or re-upload the file. |
parameter_invalid ... model_name |
Index creation only accepts marengo3.0 + pegasus1.2. Check MARENGO_MODEL/PEGASUS_MODEL. |
| Ingest succeeds but 0 segments / empty entities | The clip has little to describe (e.g. a static cartoon frame). Try a richer/longer clip. |
Health shows degraded, neo4j:false |
Neo4j unreachable — check NEO4J_URI/USERNAME/PASSWORD; make test-connection. |
Nodes tagged with the wrong domain |
DOMAIN_ID in .env differs from when you ingested. Keep it consistent; re-seed or re-tag. |
| Vector search returns nothing | Vector index not built (no embeddings were produced). Re-run make seed. |
The repo vendors a sample clip at data/videos/bbb_1080p_30fps_normal_85sec.mp4
— an 85-second excerpt of Big Buck Bunny (2008), trimmed and re-encoded
from a copy downloaded from blender.org.
- © copyright 2008, Blender Foundation / www.bigbuckbunny.org
- License: Creative Commons Attribution 3.0 (CC-BY 3.0)
Big Buck Bunny is a Creative Commons–licensed open movie: it may be reused,
redistributed, and adapted — including commercially — provided the Blender
Foundation is properly attributed. It is included here solely as a sample input
for research, testing, and demonstration of this project's pipeline; no
endorsement by the Blender Foundation is implied. See
data/videos/ATTRIBUTION.md.
Bring your own media responsibly. Any video you add is your responsibility — ensure you hold the rights or a license that permits your use. This project makes no representation about third-party clips you choose to ingest.
started from create-context-graph, repointed to video. Maintained by the community; not officially supported.