Capture browser activity · Understand it with vision AI · Visualize your day
(v0.1.0)
Chrome Extension → FastAPI → Redis/Celery → Vision API + OCR→ PostgreSQL → Dashboard
Capture is instant. AI understanding happens async, in the background.
ObserveX is a visual AI browser agent that runs entirely in your browser and on your own stack:
- A Chrome Extension (Manifest V3) silently observes meaningful browser events — navigation, tab switches, clicks, form submissions, idle resumption — and captures screenshots of the active tab.
- A FastAPI backend ingests activity events and screenshots, and enqueues them for AI analysis.
- Celery workers (backed by Redis) analyze each screenshot with Gemini Vision, falling back to PaddleOCR on any failure so no record is ever dropped.
- Structured, semantic activity metadata (activity type, application, summary, tags, confidence) is persisted to PostgreSQL.
- A React dashboard presents the day as a timeline, session list, statistics, searchable archive, and per-event detail view.
The architecture is event-driven and asynchronous: the extension only captures and uploads; all AI processing happens in background workers.
⚠️ WIP: this project is under active development. Anything not yet implemented is listed inline as TODO instead of being documented as existing.
| Area | Feature |
|---|---|
| Chrome Extension | Manifest V3 service worker, content script, popup, and options page |
| Event-driven capture | url_change, tab_switch, click, form_submit, idle_resume, and periodic interval events |
| Screenshot capture | chrome.tabs.captureVisibleTab with configurable JPEG quality |
| Offline-tolerant uploads | chrome.storage.local queue with retries (max 5 attempts) and a 1-minute flush alarm |
| Async processing | FastAPI → Redis → Celery pipeline, API never blocks on AI inference |
| Gemini Vision analysis | Structured JSON extraction (activity, application, summary, tags, confidence) |
| PaddleOCR fallback | Automatic local OCR fallback when Gemini fails (offline-capable) |
| PostgreSQL storage | Sessions, activity events, screenshots (SHA-256 dedupe hash), AI results (JSON tags/raw payload) |
| Dashboard | React 18 + Vite + Tailwind CSS |
| Timeline | Activities grouped by day |
| Session history | Session list with computed durations |
| Activity summaries | AI-generated semantic understanding per event |
| Search | Full-text LIKE search over AI summaries and activity labels |
| Statistics | Total events/sessions and top-5 activity distribution |
| Dockerized deployment | docker compose up --build runs the full stack |
flowchart LR
EXT[Chrome Extension<br/>MV3 Service Worker]
API[FastAPI API<br/>:8000]
RQ[(Redis<br/>:6379)]
CW[Celery Workers<br/>x4 concurrency]
GEM[Gemini Vision]
OCR[PaddleOCR<br/>fallback]
PG[(PostgreSQL<br/>:5432)]
DASH[React Dashboard<br/>:5173]
EXT -->|"POST /activity +<br/>POST /upload"| API
API -->|"enqueue<br/>process_screenshot"| RQ
RQ --> CW
CW -->|"1. try"| GEM
GEM -->|"any failure"| OCR
GEM -->|"AIResult"| CW
OCR -->|"AIResult"| CW
CW -->|"persist AIResult"| PG
API -->|"create event / attach screenshot"| PG
PG --> DASH
API -->|"serve screenshots<br/>/storage/*"| DASH
sequenceDiagram
participant EXT as Chrome Extension
participant API as FastAPI
participant RQ as Redis Queue
participant CW as Celery Worker
participant AI as VisionService
participant G as Gemini Vision
participant O as PaddleOCR
participant DB as PostgreSQL
EXT->>API: POST /activity (event metadata)
API->>DB: insert activity_events row
API-->>EXT: ActivityEventOut { id, ... }
EXT->>API: POST /upload?activity_id= (image bytes)
API->>API: validate type/size, store to disk, SHA-256
API->>DB: insert screenshots row
API->>RQ: enqueue process_screenshot(activity_id, path)
API-->>EXT: UploadAck { activity_id, task_id, status: "queued" }
RQ->>CW: dequeue task
CW->>AI: analyze(image_bytes)
AI->>G: generate_content(prompt + image)
alt Gemini succeeds
G-->>AI: JSON payload
AI->>AI: _extract_json (fences/prose/truncation-safe)
else Gemini fails (timeout, quota, malformed JSON)
G-->>AI: raises
AI->>O: PaddleOCR.ocr(image)
O-->>AI: extracted text (first 500 chars)
end
AI-->>CW: (AIResultSchema, source)
CW->>DB: insert ai_results row (1:1 with activity)
CW-->>RQ: mark task done (result backend)
flowchart TD
CS["Content Script<br/>click (2s debounce) / form_submit"] -->|runtime.sendMessage| BG[Background Service Worker]
T1[tabs.onUpdated → url_change] --> BG
T2[tabs.onActivated → tab_switch] --> BG
T3[idle.onStateChanged active → idle_resume] --> BG
AL[alarms → interval snapshot] --> BG
BG --> CHECK{enabled? active tab?<br/>not chrome:// URL?}
CHECK -- no --> DONE[skip]
CHECK -- yes --> CAP[captureVisibleTab<br/>JPEG, quality 60%]
CAP --> Q[(chrome.storage.local<br/>upload queue)]
Q --> FLUSH[flushQueue]
FLUSH -->|POST /activity| API
FLUSH -->|POST /upload| API
FLUSH -->|failure < 5 attempts| Q
FLUSH -->|success or 5 attempts| REMOVE[drop from queue]
| Component | Choice | Purpose |
|---|---|---|
| Language | Python 3.12 | |
| Web framework | FastAPI 0.115 + Uvicorn 0.30 | Async REST API, automatic OpenAPI docs |
| ORM | SQLAlchemy 2.0 + psycopg 3 | Typed ORM over PostgreSQL |
| Migrations | Alembic 1.13 | Schema versioning |
| Task queue | Celery 5.4 + Redis 5.0 | Distributed async workers |
| Vision AI | google-generativeai 0.8 | Gemini Vision (structured JSON output) |
| OCR fallback | PaddleOCR 2.8 + PaddlePaddle 2.6 | Local text extraction fallback |
| Validation | Pydantic 2 + pydantic-settings | Request/response schemas, env config |
| Tests | pytest + httpx | Health & validation tests |
| Component | Choice |
|---|---|
| React 18 + TypeScript 5.6 | UI |
| Vite 5 | Dev server & bundling |
| Tailwind CSS 3 | Styling |
| React Router 6 | Routing |
| TanStack Query 5 | Server-state / caching |
| Axios | HTTP client |
| Component | Choice |
|---|---|
| Chrome Manifest V3 | API surface (tabs, activeTab, scripting, storage, idle, alarms) |
| TypeScript 5.6 + esbuild 0.23 | Source & bundling (target chrome116) |
| Component | Choice |
|---|---|
| PostgreSQL 16 (alpine) | Primary database |
| Redis 7 (alpine) | Celery broker + result backend |
| Docker Compose | Single-command local stack |
ObserveX/
├── docker-compose.yml # postgres + redis + server + worker + client
├── extension/ # Chrome Extension (Manifest V3)
│ ├── manifest.json
│ ├── build.mjs # esbuild bundler (out: extension root)
│ └── src/
│ ├── background/ # service worker: capture, queue, flush, alarms
│ ├── content/ # content script: forwards click / form_submit
│ ├── lib/ # apiClient, storageManager, shared types
│ ├── popup/ # popup UI: enable/disable + queue status
│ └── options/ # options page: API URL, interval, quality
├── server/ # FastAPI backend
│ ├── Dockerfile
│ ├── requirements.txt
│ ├── alembic.ini
│ ├── alembic/ # migration environment + versions/
│ ├── tests/ # pytest suite
│ └── app/
│ ├── main.py # app factory, CORS, middleware, exception handlers
│ ├── api/routers/ # health, upload, activity, session
│ ├── config/settings.py # pydantic-settings env config
│ ├── core/ # database engine, logging, exceptions
│ ├── models/models.py # Session, ActivityEvent, Screenshot, AIResult
│ ├── repositories/ # ActivityRepository (data access layer)
│ ├── schemas/schemas.py # Pydantic request/response models
│ ├── services/ # UploadService, VisionService (Gemini + OCR)
│ ├── utils/ # (reserved)
│ └── workers/ # celery_app + process_screenshot task
└── client/ # React dashboard
├── Dockerfile
├── vite.config.ts
└── src/
├── api/client.ts # Axios API wrapper
├── hooks/queries.ts # TanStack Query hooks
├── types/index.ts # shared API types
├── components/ # Layout, ActivityCard
└── pages/ # Dashboard, Timeline, Sessions, Statistics,
# Search, Settings, ActivityDetail
- Docker + Docker Compose (easiest path) — or alternatively:
- Python 3.12 and a virtual environment
- Node.js 20+ and npm
- Google AI API key (for Gemini Vision; the PaddleOCR fallback works without it, but everything will be tagged as OCR fallback)
Copy server/.env.example to server/.env and fill in your values:
cp server/.env.example server/.env| Variable | Default | Description |
|---|---|---|
APP_NAME |
ObserveX |
Application name used in logs |
ENV |
development |
Runtime environment label |
DATABASE_URL |
postgresql+psycopg://observex:observex@postgres:5432/observex |
SQLAlchemy connection string (psycopg3 driver) |
REDIS_URL |
redis://redis:6379/0 |
Celery broker + result backend URL |
GEMINI_API_KEY |
your_gemini_api_key | Google AI Studio API key for Gemini Vision |
GEMINI_MODEL |
gemini-3.6-flash |
Gemini model identifier used for vision analysis |
MAX_UPLOAD_MB |
5 |
Maximum accepted screenshot size (MB) |
JWT_SECRET |
change-me |
TODO: reserved for future auth (no auth is enforced yet) |
Frontend (optional override at build/dev time):
| Variable | Default | Description |
|---|---|---|
VITE_API_BASE_URL |
http://localhost:8000 |
Backend base URL consumed by the dashboard (Axios) |
⚠️ Note:server/.envis git-ignored. The Docker Compose file loads it viaenv_file, anddocker-compose.ymloverridesDATABASE_URL/REDIS_URLwith the in-network service names.
docker compose up --buildThis starts five services on a shared Compose network (services reach each other by name):
| Service | Image / Build | Port | Notes |
|---|---|---|---|
postgres |
postgres:16-alpine |
5432 |
Credentials observex/observex, DB observex; healthcheck via pg_isready |
redis |
redis:7-alpine |
6379 |
Broker + result backend; healthcheck via redis-cli ping |
server |
./server |
8000 |
Uvicorn API (app.main:app) |
worker |
./server |
— | celery -A app.workers.celery_app worker --loglevel=info |
client |
./client |
5173 |
Vite dev server (React dashboard) |
Volumes
| Volume | Mount | Purpose |
|---|---|---|
observex_pg_data |
/var/lib/postgresql/data |
Persistent PostgreSQL data |
observex_screenshots |
/app/storage (server and worker) |
Shared screenshot storage; API writes, workers read |
⚠️ Note:serverandworkermust share the screenshot volume — the API stores images to disk, while the Celery worker reads them by path.
cd server
python -m venv venv
# Windows: .\venv\Scripts\activate | Linux/macOS: source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # then set GEMINI_API_KEYdocker compose up -d postgres redisalembic upgrade headuvicorn app.main:app --reloadcelery -A app.workers.celery_app worker --loglevel=infocd client
npm install
npm run devcd extension
npm install
npm run build # bundles background/, content/, popup/, options/ into extension rootThen:
- Open
chrome://extensions - Enable Developer mode
- Click Load unpacked and select the
extension/folder
⚠️ Note: the extension requires Chrome 116+ (MV3 module service worker). The default API base URL ishttp://localhost:8000— adjust it on the extension's options page if your backend runs elsewhere.
| What | Command |
|---|---|
| Full stack (Docker) | docker compose up --build |
| Backend only | uvicorn app.main:app --reload (from server/) |
| Celery worker | celery -A app.workers.celery_app worker --loglevel=info |
| Frontend dev server | npm run dev (from client/) |
| Redis | docker compose up -d redis |
| PostgreSQL | docker compose up -d postgres |
| Extension build | npm run build (from extension/) |
| Extension watch | npm run watch |
| Tests | pytest (from server/, venv active) |
| API docs (Swagger) | http://localhost:8000/docs |
| Dashboard | http://localhost:5173 |
# create a new revision after model changes
alembic revision --autogenerate -m "describe change"
# apply migrations
alembic upgrade head
# roll back one revision
alembic downgrade -1The current schema (initial revision 6a30b04d75f8):
erDiagram
SESSIONS ||--o{ ACTIVITY_EVENTS : contains
ACTIVITY_EVENTS ||--o| SCREENSHOTS : has
ACTIVITY_EVENTS ||--o| AI_RESULTS : enriched_by
SESSIONS {
string id PK
string user_id "indexed"
datetime start_time
datetime end_time "nullable"
}
ACTIVITY_EVENTS {
string id PK
string session_id FK "indexed"
datetime timestamp
text url
text page_title
string event_type "url_change|tab_switch|click|idle_resume|interval|form_submit"
}
SCREENSHOTS {
string id PK
string activity_id FK "unique"
string image_path
string hash "sha256, indexed"
datetime created_at
}
AI_RESULTS {
string id PK
string activity_id FK "unique"
string activity "e.g. coding"
string application "e.g. GitHub"
text page_title
text summary
json tags
float confidence
string source "gemini|paddleocr"
json raw_json
datetime created_at
}
Base URL: http://localhost:8000 — interactive docs at /docs.
| Method | Endpoint | Description | Payload | Response |
|---|---|---|---|---|
GET |
/health |
Liveness check | — | HealthOut |
POST |
/activity |
Create an activity event (called by the extension) | ActivityEventCreate (JSON) |
ActivityEventOut |
GET |
/activities |
List events, newest first | Query: limit (≤200, default 50), offset (≥0) |
list[ActivityEventOut] |
GET |
/activity/{activity_id} |
Event detail incl. AI result and screenshot URL | Path: activity_id |
{ event, ai_result, screenshot } |
GET |
/search |
Search AI results by summary/activity (case-insensitive LIKE) |
Query: q (required) |
list[AIResultOut] |
GET |
/sessions |
List sessions, newest first | — | list[SessionOut] |
GET |
/stats |
Aggregated counts and top-5 activities | — | StatsOut |
POST |
/upload |
Upload a screenshot; queues the process_screenshot Celery task |
Multipart: file; Query: activity_id |
UploadAck |
GET |
/storage/* |
Static serving of stored screenshots | — | image bytes |
ActivityEventCreate — body for POST /activity
| Field | Type | Notes |
|---|---|---|
session_id |
string | Created client-side (UUID); row upserted if unknown |
url |
string | Page URL |
page_title |
string | Optional, default "" |
event_type |
enum | url_change · tab_switch · click · idle_resume · interval · form_submit |
timestamp |
datetime | Optional; defaults to DB now() |
UploadAck — response for POST /upload
| Field | Type | Notes |
|---|---|---|
activity_id |
string | Linked event |
task_id |
string | Celery task ID (result backend) |
status |
string | Always "queued" |
AIResultOut — returned by /activity/{id}.ai_result and /search
| Field | Type | Notes |
|---|---|---|
id |
string | AI result row ID |
activity_id |
string | Linked event ID |
activity |
string | e.g. coding, reading |
application |
string | e.g. GitHub, YouTube |
page_title |
string | Detected title |
summary |
string | Semantic summary |
tags |
list[str] | Detected tags |
confidence |
float | 0.0 – 1.0 |
source |
string | gemini or paddleocr |
created_at |
datetime |
Error format — all errors return JSON: { "error": "<message>" }
| Status | When |
|---|---|
404 |
activity not found |
413 |
file exceeds MAX_UPLOAD_MB |
415 |
unsupported content type (allowed: image/jpeg, image/png, image/webp) |
422 |
validation failure (e.g. missing q on /search) |
500 |
unhandled exception (logged) |
The extension is intentionally thin: capture and upload only. All intelligence lives in the backend.
Event sources (background service worker):
| Event | Listener | Emitted event_type |
|---|---|---|
| Page navigation completes | chrome.tabs.onUpdated (status complete, URL changed) |
url_change |
| Active tab changes | chrome.tabs.onActivated |
tab_switch |
| Click (debounced 2 s) | content script → OBSERVEX_ACTIVITY_EVENT message |
click |
| Form submitted | content script → message | form_submit |
| Idle → active transition | chrome.idle.onStateChanged |
idle_resume |
| Periodic snapshot | chrome.alarms (observex-interval-capture, default 60 s) |
interval |
Capture & queue:
recordEvent()checks: extension enabled, tab active, URL notchrome:///chrome-extension://.- Captures the visible tab as JPEG with the configured quality (
chrome.tabs.captureVisibleTab). - Builds an
ActivityPayload(session id fromchrome.storage.local, URL, title, event type, timestamp). - Enqueues
{ payload, screenshotDataUrl, attempts: 0 }inchrome.storage.local. - Best-effort immediate
flushQueue(); a 1-minute alarm (observex-flush-queue) retries failures.
Upload queue (flushQueue) — for each queued item:
POST /activity→ create event, get itsid.POST /upload?activity_id=<id>→ upload screenshot blob.- Success → drop from queue. Failure →
attempts += 1; retried until 5 attempts, then dropped with a warning.
This makes capture resilient to network hiccups and backend restarts.
Popup shows active/paused state + pending queue count. Options page configures apiBaseUrl, captureIntervalSeconds, jpegQuality.
Implemented as the VisionService facade (server/app/services/vision_service.py), used only by the Celery worker:
flowchart LR
IMG[image_bytes] --> VS{VisionService.analyze}
VS -->|primary| GV[GeminiVisionBackend]
GV -->|"generate_content(prompt, image)<br/>temp 0.2 · max 1024 tokens"| RAW[model text]
RAW --> EX[_extract_json]
EX -->|dict| SCHEMA[AIResultSchema]
GV -->|ANY exception| OCR[PaddleOCRBackend]
OCR -->|"OCR text ≤ 500 chars"| FALLBACK[AIResultSchema<br/>activity=unknown · conf 0.3 · tags=ocr-fallback]
SCHEMA --> OUT[(source: gemini)]
FALLBACK --> OUT2[(source: paddleocr)]
- Gemini Vision is tried first. The model receives a fixed JSON-only prompt plus the raw image, with
temperature=0.2andmax_output_tokens=1024. _extract_jsonrobustly parses the model output: strips markdown code fences, tolerates surrounding prose, and recovers from token-limit truncation by trying progressively shorter JSON prefixes (up to 128 chars back).- On ANY failure (missing API key, timeout, quota, malformed JSON, empty output) the
VisionServicecatches it and falls back to PaddleOCR, which extracts visible text (truncated to 500 chars) and emits a low-confidence (0.3) result taggedocr-fallback— so the pipeline never drops a record. - The worker persists an
AIResultrow (1:1 with the activity event) including the fullraw_jsonpayload.
- Redis (
redis://redis:6379/0) acts as both the Celery broker (task transport) and result backend. - The API never performs AI work synchronously —
POST /uploadreturnsUploadAckwithstatus: "queued"immediately afterprocess_screenshot.delay(). - Queue name: default
celery(not custom-named in code).
Worker config (server/app/workers/celery_app.py):
| Setting | Value | Effect |
|---|---|---|
task_serializer / result_serializer |
json |
JSON wire format |
task_track_started |
True |
Worker started state visible in results |
task_acks_late |
True |
Acknowledges only after completion → at-least-once delivery |
worker_prefetch_multiplier |
1 |
One task at a time per worker process |
task_time_limit |
300 s |
Hard kill (soft: 270 s) |
worker_concurrency |
4 |
4 processes per worker node |
Task: process_screenshot (server/app/workers/tasks.py)
@celery_app.task(name="process_screenshot", bind=True, max_retries=3, default_retry_delay=5)
def process_screenshot(self, activity_id: str, image_path: str) -> dict:- Opens a fresh DB session (
SessionLocal). - Reads the screenshot bytes from the shared volume.
- Lazily builds (once per worker process) the
VisionService— heavy PaddleOCR import is deferred until actually needed. analyze()→(AIResultSchema, source)→ persistsAIResultrow.- On any exception →
self.retry(exc=exc)— up to 3 retries with a 5-second backoff. - Returns
{ activity_id, source, activity }to the result backend.
| Decision | Rationale |
|---|---|
| FastAPI | Async-first framework with Pydantic validation, typed schemas shared between routers/repositories, and auto-generated OpenAPI docs (/docs). |
| Redis | Lightweight in-memory broker with sub-millisecond enqueue; doubles as the result backend, avoiding an extra moving part. |
| Celery | Mature distributed task queue with retries, time limits, acks-late semantics, and horizontal scaling — exactly what a slow AI pipeline needs. |
| PostgreSQL | Relational integrity (1:1 screenshot/AI-result links, indexed FKs), plus JSON columns for flexible AI payloads (tags, raw_json). |
| Event-driven architecture | Capturing is cheap and latency-sensitive; AI inference is expensive and slow. Decoupling via a queue keeps the extension and API instant and never drops work. |
| Async workers | Gemini calls and PaddleOCR inference each take seconds — blocking the API thread would make uploads unusable. |
| Gemini Vision | Zero-shot visual understanding directly from a screenshot; structured JSON output drives the dashboard without hand-written heuristics. |
| PaddleOCR fallback | Keeps the pipeline complete even when Gemini is down, out of quota, or unconfigured — at the cost of lower confidence (0.3), explicitly surfaced via the source field. |
| Lever | How ObserveX supports it |
|---|---|
| Horizontal workers | Stateless workers: docker compose up --scale worker=4 (concurrency is per-node, 4). Tasks are re-enqueued on failure via retries. |
| Redis queue | Natural buffering — bursts of captured events are drained at worker pace; task_acks_late + prefetch 1 prevents lost work on crash. |
| Stateless API | All state lives in PostgreSQL/Redis/disk volume; API instances can be replicated behind a load balancer. |
| AI abstraction | VisionBackend interface makes the primary model or fallback swappable (e.g. another Gemini model, Claude, or a local model) without touching the pipeline. |
| Worker isolation | Heavy PaddlePaddle import is lazy and per-process; CPU-bound OCR runs only on worker nodes, not the API. |
| Database indexing | Indexes on sessions.user_id, activity_events.session_id, screenshots.hash, plus unique constraints on the 1:1 links; FKs keep joins cheap. |
| Screenshot dedupe | SHA-256 hashing is stored per screenshot — TODO: deduplication/compaction job not yet implemented. |
| What | Where |
|---|---|
| Backend env | server/.env (see Environment Variables) |
| Frontend API URL | VITE_API_BASE_URL env var, overridable in client/src/api/client.ts and the Settings page |
| Extension settings | Options page (API URL, capture interval in seconds, JPEG quality 0–1); defaults in extension/src/lib/types.ts |
| CORS origins | ALLOWED_ORIGINS setting (default ["http://localhost:5173", "chrome-extension://*"]) |
| Celery tuning | server/app/workers/celery_app.py |
# 1. Start infra (Docker)
docker compose up -d postgres redis
# 2. Backend (terminal 1)
cd server && .\venv\Scripts\activate # or source venv/bin/activate
alembic upgrade head
uvicorn app.main:app --reload
# 3. Worker (terminal 2)
cd server && celery -A app.workers.celery_app worker --loglevel=info
# 4. Frontend (terminal 3)
cd client && npm run dev
# 5. Extension (terminal 4) — then Load unpacked at chrome://extensions
cd extension && npm run watch




