A prototype pipeline that produces styled video captions from one internal Canonical Video Report (CVR) per video. Vision understanding is performed exactly once per video; each requested caption style is then generated from the CVR alone, isolating factuality from stylistic phrasing.
Supported caption styles (exhaustive): Formal, Sarcastic, Humorous-Tech, Humorous-Non-Tech.
For the full design see DESIGN.md, video_captioning_agent_spec.md, and TASKS.md.
For each task in /input/tasks.json, the agent:
- Reads and structurally validates the task entry (
task_id,video_url,styles). - Filters requested styles down to the supported set.
- Downloads the video into a per-run temporary directory (failure-isolated per task).
- Inspects the downloaded media with OpenCV (readability check + metadata).
- Samples up to 16 chronological frames at a 768px max dimension, with a 1 fps sequential fallback when frame-count metadata is unavailable.
- Calls Fireworks
Qwen2.5-VL 32B Instruct(vision mode,temperature=0.1) exactly once to produce a CVR. - Parses and validates the CVR against a strict five-field schema (
scene,primary_subjects,important_objects,timeline,overall_summary). - For each supported requested style, calls the same model in text-only mode (
temperature=0.2) sending only the serialized CVR. - Validates captions (non-empty, requested style) and logs a warning if a caption exceeds ~100 words.
- Writes
/output/results.jsonatomically as a top-level JSON array.
Task-level failures do not terminate the batch: failed tasks yield empty strings for each of their requested styles, and the pipeline always exits with code 0 after writing a valid results.json.
[
{
"task_id": "v1",
"video_url": "https://example.com/clip.mp4",
"styles": ["Formal", "Sarcastic", "Humorous-Tech", "Humorous-Non-Tech"]
}
]task_id: non-empty string. Duplicate IDs cause all tasks with that ID to be skipped.video_url: non-empty string. Invalid URLs and timeouts are treated as per-task failures.styles: list of non-empty strings. Unsupported styles are defensively ignored (not expected in valid input).
[
{
"task_id": "v1",
"captions": {
"formal": "...",
"sarcastic": "...",
"humorous_tech": "...",
"humorous_non_tech": "..."
}
}
]Each result object contains only the styles the task requested, using lowercase/snake_case keys. A missing or failed caption is represented by an empty string, never a missing key. The /output directory is created if it does not exist.
- Python 3.10+ (Python 3.12 in the Docker image)
- A Fireworks account with access to
Qwen2.5-VL 32B Instruct FIREWORKS_API_KEYset in the environment
Python dependencies (see requirements.txt):
requests>=2.31,<3
opencv-python-headless>=4.10,<5
Pillow>=10,<12
Optional local dev dependencies: pytest for the test suite.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install pytest # only for running the test suite
export FIREWORKS_API_KEY=...The entrypoint expects the POSIX paths /input/tasks.json and /output/results.json. The simplest way to provide them is to mount bind paths in Docker (see below) or, inside a Linux environment where those paths are writable:
PYTHONPATH=src python -m video_captioning_agent.pipelineTo run with custom paths from Python (e.g. for ad-hoc local runs), call run_pipeline():
from pathlib import Path
from video_captioning_agent.pipeline import run_pipeline
run_pipeline(
input_path=Path("./my_tasks.json"),
output_path=Path("./my_results.json"),
)fireworks_test_deployment.py provides two lightweight health checks against the deployed VLM:
# Verify Fireworks authentication and deployment health.
python fireworks_test_deployment.py --mode image-sanity
# Test frame sampling + video understanding against a local clip.
python fireworks_test_deployment.py --mode video --video-path /path/to/clip.mp4 --num-frames 8The Dockerfile builds a slim Python 3.12 image that copies only the runtime sources under src/ and installs requirements.txt. The entrypoint is:
ENTRYPOINT ["python", "-m", "video_captioning_agent.pipeline"]
Build and run with mounted /input / /output volumes:
docker build -t video-captioning-agent .
mkdir -p ./input ./output
cp my_tasks.json ./input/tasks.json
docker run --rm \
-e FIREWORKS_API_KEY=$FIREWORKS_API_KEY \
-v "$(pwd)/input:/input:ro" \
-v "$(pwd)/output:/output" \
video-captioning-agentResults are written to ./output/results.json. The container always exits with code 0 as long as results.json can be written. The image is well under the 10 GB compressed-image cap described in the spec.
The test suite (pytest, mocked network/model calls + small fixture media) lives under tests/. From the repo root:
PYTHONPATH=src pytest -qNo real FIREWORKS_API_KEY or large dataset is required. Each test file corresponds to a single pipeline stage (e.g. test_cvr_client.py, test_frame_sampler.py, test_pipeline.py); test_pipeline.py exercises the full mocked end-to-end flow including per-task failure isolation.
The experiments harness has its own test suite under experiments/tests/. Run both suites together from the repo root:
PYTHONPATH=src:experiments pytest tests/ experiments/tests/ -qA lightweight MLflow-backed harness lives under experiments/ for iterating on CVR prompts, frame-sampling parameters, and model/generation config. It is dev-only tooling: it is not part of the production Docker image, does not touch /input→/output, and reuses the real frame sampler (Task 6), VLM client (Task 7), and style generator (Task 10) from src/. Both CVR-generation and style-generation experiments are tracked in the same MLflow experiment (video_captioning_cvr_experiments). See EXPERIMENT_TRACKING.md for the full spec.
pip install -r experiments/requirements.txt # mlflow, pyyamlexport FIREWORKS_API_KEY=...
python experiments/run_experiment.py --config experiments/configs/exp_example.yamlOptional CLI overrides (avoid editing YAML for one-off variations):
python experiments/run_experiment.py --config experiments/configs/exp_example.yaml \
--video-path /other/clip.mp4 --num-frames 8 --max-resolution 512python experiments/run_all.pyThis loops sequentially over every *.yaml/*.yml file in experiments/configs/, logging each as its own MLflow run. A failing config logs the error and continues to the next.
Tracking metadata lives in a local SQLite database; artifacts (logged prompts and CVR JSON) live as files under ./experiments/mlruns. No mlflow server background process is required — mlflow ui with the SQLite backend is sufficient for this project's sequential, single-user use case:
mlflow ui --backend-store-uri sqlite:///experiments/mlflow.db --default-artifact-root ./experiments/mlrunsOpen http://localhost:5000, select the video_captioning_cvr_experiments experiment, and use the runs table (one row per run, params/metrics as columns) and the Compare view to diff configurations side by side. Each run's system_prompt.txt, user_prompt_template.txt, cvr_output.json, and style_captions.json are viewable as artifacts.
Each YAML config supports: name, video_path, system_prompt (or system_prompt_path), user_prompt_template (or user_prompt_template_path), frame_sampling (num_frames, max_resolution, fallback_fps), cvr_model (name, temperature, max_tokens), and style_model (name, temperature, max_tokens). See experiments/configs/exp_example.yaml for a fully annotated example. Omitted sections default to the production constants from src/.
.
├── src/video_captioning_agent/
│ ├── contracts.py # VideoTask, CVR, FrameSample, VideoMetadata, TaskResult
│ ├── input_loader.py # /input/tasks.json parsing + structural validation
│ ├── styles.py # Supported style set + task eligibility (input validation)
│ ├── downloader.py # Bounded, per-task failure-isolated downloads
│ ├── video_inspection.py # OpenCV readability + metadata extraction
│ ├── frame_sampler.py # Uniform sampling + sequential 1 fps fallback
│ ├── cvr_client.py # CVR prompt construction + Fireworks VLM client (config-swappable)
│ ├── cvr_parser.py # Strict JSON/CVR validation, no fact fabrication
│ ├── style_generator.py # CVR-only text requests, all 4 styles unconditionally (gpt-oss-120b)
│ ├── caption_validation.py # Validates all 4 captions; ~100-word warning-only
│ ├── result_writer.py # Atomic /output/results.json + requested-style filtering
│ └── pipeline.py # Sequential orchestration with failure isolation
├── tests/ # pytest suite, one file per stage
├── experiments/ # Dev-only MLflow experiment harness (not in Docker image)
│ ├── run_experiment.py # Single-run CLI
│ ├── run_all.py # Batch runner over configs/
│ ├── harness.py # Config loading + MLflow setup
│ ├── configs/ # YAML experiment configs
│ └── tests/ # Experiments test suite
├── Dockerfile # Runtime container image (src/ only)
├── requirements.txt # Runtime dependencies only
├── fireworks_test_deployment.py # Optional manual Fireworks sanity script
└── AGENTS.md / DESIGN.md / video_captioning_agent_spec.md / TASKS.md / EXPERIMENT_TRACKING.md
- Audio is out of scope; captioning is visual-only.
- No programmatic caption factuality verifier in this pass — the primary safeguard is CVR-only input isolation at the style-generation stage (see
TASKS.md→ Future Scope). - Frame padding is intentionally disabled: short or degraded videos return only the unique frames that could actually be read.
- Never commit
.env,FIREWORKS_API_KEY, signed upload URLs, model weights, or large dataset files. The Fireworks API key is read from the environment, never hardcoded.