Built for Hack Devengers 1.0. Analyzes a recorded or live interview clip across seven independent, fully-local signals and fuses them into an evidence-backed trust report for a human reviewer — it never makes an automated accept/reject decision on its own.
| Frontend | veritas-atifa.netlify.app |
| Backend API | veritas-backend-drhzf2c8c3gafecs.centralindia-01.azurewebsites.net |
| Health check | /api/health → {"status":"ok","service":"veritas-api"} |
| API docs (Swagger) | /docs |
| OpenAPI spec | /openapi.json |
Architecture:
Netlify (React/Vite frontend)
│
▼
Azure App Service (FastAPI + Gunicorn/Uvicorn backend)
│
▼
Veritas signal-analysis pipeline
- Netlify is connected directly to the GitHub repo — pushing to GitHub triggers an automatic Netlify redeploy of the frontend.
- The frontend's API base URL (
frontend/src/api.js) points at the Azure backend in production:https://veritas-backend-drhzf2c8c3gafecs.centralindia-01.azurewebsites.net/api(previously/api, a local dev proxy path — see Setup below for local development, which still uses the proxy). - The backend root
/intentionally returns{"detail":"Not Found"}— there is no root route defined; use/api/healthto check liveness.
Remote hiring interviews are increasingly targeted by deepfake/face-swap candidate fraud (the FBI issued a PSA on this exact pattern), screen-replay attacks (holding a phone/tablet playing footage up to the webcam — cheaper and more common than a real-time deepfake), identity swaps mid-session, and live coaching via a second monitor or earpiece. Existing deepfake detectors are single pretrained classifiers trained on specific generation methods — they generalize poorly to unseen tools, which is exactly the failure mode that sinks them in the real world. Veritas doesn't trust any single signal. It combines seven independently-computed, explainable signals spanning three distinct threat classes and only escalates when several agree.
| Signal | Threat class | What it measures | How |
|---|---|---|---|
| Visual authenticity | Synthetic media | Frequency-domain + noise-residual synthesis artifacts | FFT high-frequency energy ratio + face/background noise-residual comparison (classical forensics, no model download) |
| Audio-visual sync | Synthetic media | Whether lip movement tracks the audio | Cross-correlation of mouth-openness (mediapipe landmarks) against the audio RMS envelope |
| Liveness & behavior | Synthetic media | Natural human micro-behavior | Blink rate/regularity (EAR) + head-pose micro-jitter (solvePnP) |
| Active challenge-response | Synthetic media | Real-time reaction to an instruction the subject couldn't have pre-rendered | A random motion challenge (turn head / nod / shake / blink sequence) issued mid-recording, verified from landmark motion — no ASR dependency |
| Screen / replay-attack detection | Replay attack | Whether the "face" is actually a screen held up to the camera | Moire-pattern FFT peak detection + specular-glare geometry + Hough-line bezel/edge detection + refresh-rate luminance flicker |
| Identity continuity | Identity swap | Whether the same person is on camera throughout the session | Pose-normalized geometric landmark ratios (interocular, jaw, mouth, nose) compared against an early-session baseline via robust (MAD-based) z-scoring |
| Off-screen gaze / coaching indicators | Live coaching | Sustained gaze away from the screen not compensated by a head turn | Iris-landmark offset (mediapipe refine_landmarks) vs. head yaw (solvePnP) |
All seven scores + confidences are fused in fusion.py into one
overall_score, overall_status, and a timestamped evidence ledger. No
score in this system is ever random or placeholder — every number traces to
a measured quantity.
- Backend: FastAPI + OpenCV + mediapipe (pinned to
0.10.13for the bundled-model legacy API — no network calls, no API keys) + librosa + ffmpeg (system binary, for audio extraction). Served via Gunicorn/Uvicorn workers in production (Azure App Service). - Frontend: React + Vite, plain CSS design system (no Tailwind — kept dependency-light and dead simple to run). Deployed to Netlify.
- 100% free/open-source. No paid APIs, no API keys, nothing that expires.
Requires Python 3.10+ and ffmpeg on PATH.
# Windows: install ffmpeg via https://ffmpeg.org/download.html (or `choco install ffmpeg`)
# and make sure it's on PATH — check with `ffmpeg -version`
cd backend
python -m venv venv
venv\Scripts\activate # Windows PowerShell
# source venv/bin/activate # macOS/Linux
pip install -r requirements.txt
uvicorn app.main:app --reload --port 8000Backend is now live at http://127.0.0.1:8000. Check http://127.0.0.1:8000/api/health.
In a second terminal:
cd frontend
npm install
npm run devOpen http://127.0.0.1:5173. In local dev, src/api.js uses the /api
proxy path and the Vite dev server proxies all /api/* requests to the
backend on port 8000 — no CORS setup needed. In production, src/api.js
points directly at the deployed Azure backend URL (see
Live deployment above).
- Start a session → grant camera access.
- Hit Start 12s recording. ~2.5s in, a random live challenge appears on screen (e.g. "Turn your head to the left, then back to center") — this is the hardest signal to spoof because it's issued during capture.
- Recording auto-submits to
/api/analyzewhen it ends. - The pipeline runs (frame extraction → 7 signals → fusion) and returns a Trust Report: overall score/status, per-signal breakdown with confidence, and a timestamped evidence ledger.
- Use the human review panel to approve / flag / reject and attach a note — this is logged against the session, demonstrating the human-in-the-loop design (Veritas recommends, a person decides).
For a judge who wants to try to fool it live: have them hold up a phone running a face filter/deepfake app to the webcam mid-demo, or use the file upload path with a pre-recorded manipulated clip — the "or analyze an existing file" option under the recorder supports this without needing a live camera.
veritas/
├── backend/
│ ├── requirements.txt
│ └── app/
│ ├── main.py # FastAPI routes, orchestrates the pipeline
│ ├── config.py # env-driven settings
│ ├── storage.py # in-memory session/challenge store (demo-scoped)
│ ├── schemas.py # pydantic response models
│ └── pipeline/
│ ├── video_io.py # ffmpeg/opencv ingestion, frame sampling
│ ├── face_detect.py # mediapipe face box + 478-pt landmarks (incl. iris)
│ ├── visual_authenticity.py # FFT + noise-residual artifact signal
│ ├── av_sync.py # mouth-motion / audio cross-correlation
│ ├── liveness.py # blink rate + head-pose jitter
│ ├── challenge.py # active challenge issuance + verification
│ ├── screen_replay.py # moire/glare/bezel/flicker replay-attack detection
│ ├── identity_continuity.py # geometric same-person verification across session
│ ├── gaze_coaching.py # iris-based off-screen gaze / coaching indicators
│ └── fusion.py # combines all 7 signals into TrustReport
└── frontend/
├── index.html
├── vite.config.js
└── src/
├── main.jsx
├── App.jsx # view state machine: hero → capture → loading → report
├── api.js # fetch wrappers for /api/* — points at Azure in production
├── index.css # design system (tokens, components)
└── components/
├── Header.jsx
├── Hero.jsx
├── CapturePanel.jsx # webcam capture, challenge overlay, upload fallback
├── LoadingState.jsx
├── ScoreDial.jsx
├── SignalCard.jsx
├── EvidenceTimeline.jsx
├── ReviewPanel.jsx
└── TrustReport.jsx
Base URL (production): https://veritas-backend-drhzf2c8c3gafecs.centralindia-01.azurewebsites.net/api
GET /api/health— liveness checkPOST /api/challenge— issues a random live challenge{id, type, instruction, window_seconds}POST /api/analyze— multipart form:video(file), optionalchallenge_id,challenge_offset_s→ returns{report: TrustReport}POST /api/session/{id}/review—{decision: "approved"|"flagged_for_review"|"rejected", note}GET /api/session/{id}— fetch a stored report + reviewGET /api/sessions— list all sessions (demo/session-scoped, in-memory)
Full interactive docs (Swagger UI) and the raw OpenAPI schema are linked in Live deployment above.
- Storage is in-memory, scoped to one running process — swap for Postgres/SQLite for anything beyond a demo. In the current Azure deployment this means session data does not persist across backend restarts.
- AV-sync is an honest lightweight proxy (mouth-motion tracks audio
envelope), not a SyncNet-equivalent lip-reading match — documented as
such in
av_sync.pyrather than oversold. - Identity continuity uses pose-normalized landmark geometry ratios, not a trained face-recognition embedding (that would require a model download, which conflicts with the fully-local guarantee). It's a materially weaker identity signal than a biometric embedding, and says so — it answers "did the coarse facial geometry change in a way inconsistent with the same person," not "is this a verified biometric match." Its confidence is deliberately capped lower than the other signals to reflect this.
- Off-screen gaze detection cannot distinguish "reading a second monitor" from "glancing at notes on the desk" or "looking at a clock" — it reports a behavioral pattern (sustained, head-uncompensated gaze shift) for human review, not proof of coaching.
- Screen-replay detection can be partially evaded by a high-end anti-glare display and careful framing — the four cues (moire/glare/bezel/flicker) raise attacker cost, they don't guarantee detection.
- No auth on the API — fine for a hackathon demo, not for production. This applies to the live Azure deployment as well: the API is currently open, with no API keys or auth in front of it.
- Forensic signals are heuristic and affected by lighting/webcam quality — the UI and the API response both say so, and the system is designed to recommend human review, never to auto-reject a candidate.
- Swap in-memory storage for a real DB with an audit trail.
- Add auth + per-org API keys for ATS integration (Greenhouse/Lever webhook).
- Expand the evaluation harness with a labeled dataset of real vs. synthetic/replayed/swapped interview clips to validate/tune the anomaly thresholds currently set from documented heuristic ranges — this applies to all seven signals.
- Replace the identity-continuity geometric descriptor with a proper local face-embedding model (e.g. a small ONNX face-recognition network) in environments where a one-time model download is acceptable — this would meaningfully strengthen same-person verification beyond landmark ratios.
- Optional: the
try_load_pretrained_classifier()hook invisual_authenticity.pyis a ready opt-in slot for a pretrained deepfake classifier as an additional signal, for environments with model-download access. - Add evidence provenance (hash/sign the source video + evidence frames) so the Trust Report itself becomes tamper-evident — flagged in the prior architecture review as a cheap, high-value addition not yet built.