WatchPulse is a streaming discovery product that helps users decide what to watch based on their region, streaming services, and preferences. External APIs populate a local catalog; browsing and filter changes will query WatchPulse's own data rather than calling upstream APIs.
The first-release UI provides region- and provider-aware discovery across four shared, filterable sections:
- Top 10
- New Releases
- Recently Added
- Upcoming
The evidence-only Leaving Soon API remains available, but its UI section is deferred until expiration ingestion is enabled.
See AGENTS.md for the current product requirements and
docs/architecture.md for the active architecture and
implementation decisions. AGENTS.md remains the authoritative product brief.
Documentation is split by responsibility:
Version v0.5 is the current release. The repository includes:
- configurable TMDB discovery by region and provider;
- full movie and TV metadata ingestion;
- TMDB watch-provider ingestion;
- region/provider-scoped streaming lifecycle ingestion;
- append-only lifecycle history and idempotent DuckDB availability state;
- bounded manual ingestion workflow;
- append-only raw Parquet storage;
- retrying and rate-limited HTTP access;
- source-independent content and provider models;
- dbt-duckdb staging models for TMDB discovery and streaming events;
- dbt intermediate models for canonical content, lifecycle events, current availability, and upcoming availability;
- a tested local
catalog_availabilitymart with stable provider names and TMDB-preferred metadata; - catalog freshness metadata and failure-safe atomic DuckDB publication;
- normalized content, genre, provider-source, availability, and event marts;
- a read-only FastAPI backend with freshness and catalog reference endpoints;
- a parameterized global filter engine with Top 10, New Releases, Recently Added, Upcoming, and Leaving Soon endpoints;
- scoped local title details with aggregated provider availability;
- a responsive React frontend with region-aware provider selection and shared content type, genre, runtime, release-year, rating, and language filters;
- filter-reactive Top 10, New Releases, Recently Added, and Upcoming rails with reusable poster cards and explicit loading, empty, failure, and retry states;
- local title search that respects the active region, providers, and filters;
- verified provider actions plus explicit outbound TMDB title-detail links;
- bounded broad and priority-based incremental TMDB enrichment planning;
- unit tests that do not make live API calls.
Version v0.5 adds the complete local guest discovery interface, catalog
search, Upcoming lifecycle presentation, TMDB/provider navigation, and
intelligent enrichment planning. See the
v0.5 release notes.
Add STREAMING_AVAILABILITY_API_KEY only when running live lifecycle ingestion;
offline tests and warehouse builds do not require it.
- Python 3.12 recommended (Python 3.10 or newer is required)
- Node.js 22 for v0.5 frontend development (
npmis included with Node.js) - A TMDB v3 API key
pyenvandpyenv-virtualenvare optional
Get a TMDB API key from the TMDB API settings page.
Create a local pyenv environment named watchpulse if it does not already
exist. pyenv local creates an ignored .python-version file for your machine:
pyenv install 3.12.13 # skip if already installed
pyenv virtualenv 3.12.13 watchpulse # skip if already created
pyenv local watchpulse
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"Install the optional dbt/DuckDB warehouse toolchain when working on warehouse models or the local discovery API:
python -m pip install -e ".[dev,warehouse]"If you prefer a different Python 3.12 patch release, create the watchpulse
environment from that version instead.
Pyenv is not required. Any supported Python installation can be used:
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"On Windows PowerShell, activate the environment with:
.venv\Scripts\Activate.ps1Create a local environment file and add your TMDB key:
cp .env.example .envTMDB_API_KEY=your_key_hereThe remaining values in .env.example have development defaults. Important
settings include:
SUPPORTED_REGIONS: comma-separated ISO country codes, such asGR,GB;SUPPORTED_PROVIDERS: stable WatchPulse provider keys;LAKE_ROOT: location of the append-only Parquet lake;DATABASE_PATH: operational ingestion-state database path;WATCHPULSE_SERVING_DB_PATH: atomically published dbt serving database;- discovery windows for new, recently added, and leaving-soon content.
Never commit .env; it is ignored by Git.
Build and atomically publish the local serving database, then start FastAPI:
make dbt-publish
uvicorn watchpulse.api:create_app --factory --reloadOpen http://127.0.0.1:8000/docs for the interactive OpenAPI interface. The
API reads only WATCHPULSE_SERVING_DB_PATH; requests never call TMDB or the
Streaming Availability API. See the API guide for routes,
filters, section semantics, and examples.
The Python installation does not provide Node.js. If npm is not found, use
the committed frontend/.nvmrc with Node Version Manager:
cd frontend
nvm install
nvm use
cd ..You can instead install Node.js 22 through its official operating-system installer. Keep the API running, then use a second terminal:
make frontend-install
make frontend-devOpen http://127.0.0.1:5173. The v0.5 frontend reads API health, catalog
freshness, reference filters, Top 10, New Releases, Recently Added, and Upcoming
from FastAPI, proving the complete local serving path. The browser never calls
either upstream data provider. See the
frontend guide for configuration and quality commands.
Lifecycle-only titles can be enriched incrementally with canonical TMDB metadata and current watch-provider evidence without repeating broad discovery:
python -m ingestion.enrich_streaming_metadata --event-type upcoming --country GR
make dbt-publishThe command independently skips title IDs with retained metadata or provider responses and reports its bounded TMDB request count, and makes no Streaming Availability requests.
Run the complete Greece refresh for Netflix, Disney+, Prime Video, and Apple TV+:
make catalog-refreshThe guarded workflow performs complete TMDB movie/TV discovery, publishes a
validated catalog snapshot, ingests all pages of Streaming Availability new
and upcoming changes within a 100-request run budget, incrementally enriches
missing/stale TMDB metadata, then atomically publishes the final DuckDB. Broad
watch-provider enrichment is opt-in. The previous serving database remains live
if discovery, dbt validation, or candidate publication fails.
Every discovery page belongs to a run manifest. dbt selects only the latest complete run for each region/provider/content-type combination, so samples and interrupted refreshes never leak into the frontend. A title absent from the newest complete snapshot is no longer considered currently available.
Useful bounded variants:
# Current catalog only; no lifecycle or per-title enrichment.
python -m ingestion.full_refresh --country GR --skip-streaming --skip-enrichment
# Clean metadata backfill after discovery; provider payloads remain disabled.
python -m ingestion.full_refresh --country GR --enrichment-mode backfill
# Explicit exceptional provider enrichment.
python -m ingestion.full_refresh --country GR --include-watch-providers
# Refresh one provider while preserving the latest complete snapshots of the others.
python -m ingestion.full_refresh --country GR --provider netflixThe default full-refresh summary is retained locally at
data/full-refresh-summary.json. Streaming request usage is also recorded in
the operational DuckDB and constrained by the configured monthly cap.
pyproject.toml is the canonical project configuration. It contains package
metadata, supported Python versions, runtime dependencies, development extras,
and pytest/Ruff settings.
For development, install the editable project and its development tools:
python -m pip install -e ".[dev]"requirements.txt remains as a small compatibility entry point for deployment
services that require that filename. Installing it resolves the project and its
runtime dependencies from pyproject.toml, avoiding two manually synchronized
dependency lists:
python -m pip install -r requirements.txtAfter installing the development dependencies, install both repository hooks:
make install-hooksThe hooks are intentionally split by cost:
- before commit: direct commits to
main/masterare blocked, Ruff checks and formats Python, repository hygiene checks run,detect-secretsscans staged text files, and dbt is compiled with its documentation generated; - before push: direct pushes to
main/masterare blocked and the complete offline pytest suite runs with coverage enforcement; - commit message: Commitizen enforces Conventional Commits.
Secret detection uses the committed .secrets.baseline. If a legitimate test
fixture is detected, audit the finding instead of bypassing the hook:
detect-secrets audit .secrets.baselineNever add a real credential to the baseline. Rotate it immediately if it was ever staged or committed.
Run either stage manually when needed:
make precommit
make prepushRun the test suite with a terminal coverage summary, XML report, and minimum coverage gate:
make coverageInstall the pinned dbt packages once before the first dbt hook run, or whenever
warehouse/packages.yml changes:
make dbt-depsGitHub Actions does not invoke Git branch-protection hooks. It runs the same
underlying lint and coverage Make targets directly, then validates the dbt
foundation with dbt deps, dbt debug, and dbt parse against a temporary
DuckDB file. Branch protection therefore remains a local commit/push safeguard
and needs no CI exclusion.
Hooks use the active project environment, so run
python -m pip install -e ".[dev]" after dependency changes.
Create a guided conventional commit with:
cz commitExamples of accepted messages are feat: add provider mapping,
fix(ingestion): handle an empty page, and docs: update roadmap. Commitizen
uses the PEP 621 version in pyproject.toml; when a version is ready, preview
and apply the semantic bump with:
cz bump --dry-run
cz bumpThe bump updates the project version and changelog and creates a vX.Y.Z tag.
Start with a smoke test. It fetches one discovery page for every configured region, provider, and entity type. Discovery-only is the default, so this makes no per-title metadata or watch-provider requests:
python -m ingestion.run --max-pages 1Run the complete configured catalog ingestion with:
python -m ingestion.runThe summary reports expected/fetched pages and whether every provider/content query completed. Full per-title enrichment is deliberately opt-in and should only be used for a bounded set while the incremental enrichment queue is built:
python -m ingestion.run --max-pages 1 --enrichPlan the dedicated catalog enrichment locally before making any TMDB requests:
python -m ingestion.enrich_catalog --mode incremental --dry-runThe initial broad backfill selects titles that have never retained metadata or provider payloads. Set an explicit cap appropriate for the catalog size:
python -m ingestion.enrich_catalog --mode backfill --max-titles 10000 --dry-run
python -m ingestion.enrich_catalog --mode backfill --max-titles 10000Afterward, routine incremental runs deduplicate titles, prioritize product value, and refresh only new or stale payloads:
python -m ingestion.enrich_catalog --mode incrementalAdd --metadata-only to skip watch-provider payloads. Add
--plan-output data/enrichment-plan.json to save the complete local plan for
inspection. Planning never calls TMDB; only execution does. Completed batches
are retained as immutable Parquet files, allowing later plans to resume.
See the catalog refresh runbook for the twice-weekly operating flow, clean rebuild procedure, diagrams, success gates, request expectations, and recovery commands.
To override the configured regions for one run, repeat --country:
python -m ingestion.run --country GR --country GB --max-pages 1Raw responses are written beneath:
data/lake/raw/source=tmdb/endpoint=.../entity_type=.../country=.../date=...
Each batch creates an immutable Parquet file. Re-running ingestion preserves the previous raw responses; downstream transformations will deduplicate records at their documented business grain.
Inspect a readable sample from the locally stored catalog without making new API calls:
python -m ingestion.inspect --country GR --limit 20This is a raw-ingestion diagnostic over the TMDB subscription-catalog seed. It is not the normalized lifecycle-aware serving query exposed by the v0.4 API.
Before a full backfill, verify that the configured TMDB provider IDs are still valid for the selected region:
python -m ingestion.sources.tmdb.list_providersThis command and ingestion require network access and a valid TMDB_API_KEY.
The normal test suite does not.
After adding a direct Movie of the Night API key to .env, run a one-request
Greece smoke test:
python -m ingestion.run_streaming_availability \
--country GR \
--change-type new \
--max-requests 1 \
--max-pages-per-type 1A normal bounded MVP run requests one page each for new and upcoming:
python -m ingestion.run_streaming_availability --country GRThis costs at most two requests with the default configuration. The client and
event model still support removed, updated, and expiring, but those types
and the Leaving Soon section are deferred beyond the first public release.
WatchPulse cards and search results link explicitly to TMDB for richer title
details instead of duplicating a full internal details page.
Full pagination is opt-in and remains protected by the request ceiling:
python -m ingestion.run_streaming_availability \
--country GR \
--change-type new \
--all-pages \
--max-requests 500The manual GitHub workflow uses that full-pagination mode for the four combined
subscription catalogs (netflix, disney_plus, prime_video, and
apple_tv_plus). It has no cron trigger, follows both selected cursor chains to
completion, and stops before request attempt 501. Ordinary local runs remain
limited to one page per type unless --all-pages is supplied.
The configured monthly cap relies on the DuckDB usage ledger and therefore applies across runs only where that database persists. Until durable workflow storage is introduced, each manual GitHub invocation must be treated as having its own 500-request ceiling.
The runner combines all configured subscription providers per request, writes
raw pages under data/lake/raw, writes normalized append-only events under
data/lake/events, and records usage in DuckDB. It enforces both the per-run and
calendar-month request limits from .env, including retry attempts.
Inspect or replay locally persisted events:
python -m ingestion.inspect_events --country GR --limit 20
python -m ingestion.replay_events --country GRRun all tests from the repository root:
python -m pytest -qCurrent tests cover raw-lake writes, HTTP retry behavior, configuration validation, and translation from TMDB payloads into internal content models.
GitHub Actions currently provides:
CI: installs Python 3.12 dependencies, checks syntax, and runs all tests on pull requests and pushes tomaster;Manual TMDB ingestion: runs a bounded, manually triggered ingestion and uploads its raw Parquet lake as a seven-day workflow artifact.
To use manual ingestion, add TMDB_API_KEY under the repository's Actions
secrets, then choose Actions → Manual TMDB ingestion → Run workflow. The
workflow intentionally requires a page limit so an accidental manual run cannot
start an unbounded catalog backfill.
Scheduled ingestion will be enabled in v0.2, once incremental lifecycle
ingestion exists. Production continuous deployment will be added in v0.6
after the hosting and frontend decisions are recorded; adding a pretend deploy
job before a target exists would not produce a working deployment.
watchpulse/ Shared configuration and domain models
ingestion/core/ Source contracts, HTTP behavior, and raw-lake writes
ingestion/sources/tmdb/ TMDB client, adapter, and provider configuration
ingestion/tests/ Offline unit tests and fixtures
warehouse/ Planned dbt-duckdb transformation project
api/ Planned read-only discovery backend
frontend/ React guest-first discovery UI
scripts/ Planned safe operational commands
tests/ Cross-component and end-to-end tests
.github/workflows/ CI, ingestion automation, and future deployment
docs/ Architecture documentation
data/ Local generated data (gitignored)
Planned boundaries are warehouse/ for dbt and DuckDB models, api/ for the
local query layer, and frontend/ for discovery UI. The frontend will only call
the WatchPulse API; it will never call TMDB or a streaming availability source.
- Finish foundation configuration, DuckDB setup, models, and ingestion run observability.
- Add Streaming Availability API ingestion and historical lifecycle events.
- Build dbt staging, normalized models, and the serving catalog.
- Add safe parameterized discovery queries and API endpoints.
- Build the filterable frontend and content details.
- Add CI, scheduled ingestion, monitoring, and deployment.
- Add bounded natural-language discovery after deterministic discovery works.
Authentication and personalization are intentionally deferred. Core discovery will work without an account.