Production-minded fraud platform built as a modular monorepo with Kafka ingestion, a Bytewax stream worker, Redis online features, XGBoost plus rules hybrid scoring, FastAPI business APIs, PostgreSQL persistence, MLflow model registry, Evidently drift reporting, Grafana dashboards, and a real Next.js analyst console.
Fraud teams need more than an offline classifier. They need a system that can ingest live payment events, maintain rolling behavioral context, combine deterministic controls with ML scoring, surface cases to analysts quickly, and close the loop with feedback, monitoring, and retraining.
- Most fraud portfolios stop at notebooks and batch metrics.
- Real-world fraud systems are judged on latency, explainability, observability, and analyst workflow.
- This repo focuses on the full product story: streaming ingestion, online features, hybrid decisioning, case review, drift monitoring, and demo-friendly operations tooling.
|
|
- Real-time producer to Kafka to Bytewax to Redis to scoring to Postgres pipeline
- Redis-backed rolling features for velocity, spend, novelty, and geo jumps
- Hybrid decisioning with rules plus champion XGBoost model
- Persisted rule hits, reason codes, scores, and case metadata
- API-driven analyst console with live polling-based updates
- Analyst feedback loop written to both PostgreSQL and
tx.feedback - MLflow model registry plus Evidently drift reporting
- Grafana and Prometheus dashboards for throughput, latency, DLQ, and drift
flowchart LR
Producer[Producer\nsynthetic tx generator] -->|tx.raw| Kafka[(Kafka)]
Kafka --> Worker[Bytewax stream worker]
Worker -->|tx.validated / tx.enriched / tx.scored / tx.decisions| Kafka
Worker --> Redis[(Redis online feature store)]
Worker --> Postgres[(PostgreSQL)]
API[FastAPI business API] --> Postgres
API --> Redis
API --> Kafka
Trainer[Trainer / ML Ops] --> MLflow[(MLflow)]
Trainer --> Evidently[Evidently reports]
Trainer --> Postgres
Analyst[Next.js analyst console] --> API
Prometheus[(Prometheus)] --> Grafana[(Grafana)]
API --> Prometheus
Producer --> Prometheus
Worker --> Prometheus
Trainer --> Prometheus
More detail:
- Architecture overview
- Transaction lifecycle
- Runbook
- Troubleshooting
- Demo script
- Release notes + pinned text
- GitHub profile kit
| Layer | Tools |
|---|---|
| Streaming | Kafka, Bytewax |
| Online state | Redis |
| ML and rules | XGBoost, scikit-learn, YAML rule engine |
| APIs | FastAPI, Pydantic v2 |
| Persistence | PostgreSQL, SQLAlchemy, Alembic |
| Analyst UI | Next.js 15, TypeScript, Tailwind, shadcn-style components |
| Observability | Prometheus, Grafana |
| Model Ops | MLflow, Evidently |
| Packaging | Docker Compose, Makefile, GitHub Actions |
apps/
analyst-console/ # Next.js internal operations UI
api/ # FastAPI business API
producer/ # synthetic traffic generator + export CLI
stream-worker/ # Bytewax flow + stream runtime
trainer/ # XGBoost training, MLflow, Evidently
libs/
common/ # config, logging, FastAPI service helpers
contracts/ # shared Pydantic event contracts
feature_engineering/# online/offline feature computation
feature_store/ # Redis + in-memory feature store adapters
model_runtime/ # champion model loading and hybrid scoring
observability/ # Prometheus metrics
persistence/ # SQLAlchemy models and repositories
rules/ # YAML-driven fraud rules
infra/
docker/ # Dockerfiles
grafana/ # dashboards + alerts as code
kafka/ # topic bootstrap
postgres/ # init SQL + Alembic migrations
prometheus/
docs/
tests/
- Docker Desktop or Docker Engine with Compose
- Node 20+ only if you want to run the analyst console outside Docker
- Python 3.11 for local non-Docker backend development
docker compose up --buildThe stack includes:
kafkaredispostgresprometheusgrafanamlflowdb-migratemodel-bootstrapproducerstream-workerapitraineranalyst-console
Use this when you want a deterministic local reset instead of reusing old Kafka, Redis, PostgreSQL, MLflow, or Grafana state.
make reset-stack
make bootstrapEquivalent Docker Compose sequence:
docker compose down --volumes --remove-orphans
docker compose up -d --build| Surface | URL |
|---|---|
| Analyst console | http://localhost:3001 |
| Overview live demo | http://localhost:3001/overview |
| Analyst backlog live demo | http://localhost:3001/cases |
| API docs | http://localhost:8000/docs |
| Producer | http://localhost:8001/producer/status |
| Stream worker | http://localhost:8002/worker/status |
| Trainer | http://localhost:8003/training/status |
| Grafana | http://localhost:3000 |
| Prometheus | http://localhost:9090 |
| MLflow | http://localhost:5000 |
OverviewandCasesenable live mode by default.- Live mode is implemented with real polling-based browser updates, not fake frontend animation.
- Polling hits FastAPI live endpoints:
/dashboard/live/cases/live
- Turning live mode off freezes the current browser snapshot until you click refresh.
- The backlog page defaults to the real analyst queue:
decision=REVIEWandstatus=open.
- The analyst console is the fraud operations surface.
- It is where you review cases, inspect scores and reasons, trigger demo bursts, and submit feedback.
- Grafana is the operator and observability surface.
- It is where you show throughput, latency, errors, drift, and alerting.
- The analyst console is intentionally not pretending to be Grafana.
The overview page includes real producer controls that call API-backed endpoints:
POST /demo/producer/startPOST /demo/producer/stopPOST /demo/producer/burstPOST /demo/producer/boostPOST /demo/producer/reset
These controls change real backend behavior and push new events through Kafka, Bytewax, Redis, scoring, Postgres, and the UI.
- Open
http://localhost:3001/overview. - Point out the live badge, last-updated label, recent-window counters, and activity feed.
- Trigger
Inject impossible travel burstorInject new-device high-amount burst. - Open
http://localhost:3001/casesand show new rows rising into the real analyst backlog. - Open a case detail page and submit feedback.
- Open Grafana to show the operational view of the same system.
- Open MLflow to show the registered champion model.
- The trainer can bootstrap a champion XGBoost model from generated CSV data.
- It can also rebuild training data from persisted PostgreSQL transactions.
- When analyst feedback exists, retraining uses the latest analyst label instead of the synthetic source label.
- MLflow stores registered model versions and aliases like
champion. - Evidently generates drift artifacts and updates drift metrics surfaced in Grafana.
- Prometheus metrics are exposed by every Python service on
/metrics. - Grafana dashboards are provisioned from
infra/grafana/dashboards. - Alert rules are provisioned from
infra/grafana/provisioning/alerting. - Local alert notifications terminate at the API webhook sink at
POST /ops/grafana-alerts. - Drift gauges are updated after the trainer generates a new Evidently report.
docker compose up --build
make reset-stack
make bootstrap
docker compose logs -f stream-worker
docker compose logs -f api
docker compose exec trainer fraud-trainer-cli bootstrap-model --force
docker compose exec trainer python -c "from fraud_platform_common.config import RuntimeSettings; from fraud_platform_persistence import FraudRepository; repo = FraudRepository(RuntimeSettings(service_name='trainer')); rows = repo.training_frame(); print(rows[-1]['label_source'], rows[-1]['latest_feedback_label'])"
docker compose exec trainer fraud-trainer-cli drift-report --sample-size 500
docker compose exec producer fraud-producer-cli export-dataset --output /data/bootstrap_transactions.csv --events 3000
docker compose exec api python -c "import requests; print(requests.get('http://localhost:8000/dashboard/overview').json())"
curl http://localhost:8000/dashboard/live
curl "http://localhost:8000/cases/live?status=open&decision=REVIEW"
curl -X POST http://localhost:8000/demo/producer/burst -H "content-type: application/json" -d "{\"scenario\":\"impossible_travel\",\"count\":10}"
make test-docker
make frontend-testThese screenshots were captured from the running local stack and wired directly into the repo so visitors can see the product story before they clone it.
Live overview with recent-window counters, API-backed live polling, producer controls, and the latest persisted decisions.
Grafana remains linked from the monitoring page and README URLs because the local demo stack keeps the operator dashboard behind its own login surface.
See docs/screenshots-checklist.md for the ordered capture plan and suggested captions.
python -m compileall apps libs
pytestImportant: this repo targets Python 3.11. Running tests with Python 3.10 will fail because the code intentionally uses Python 3.11 features such as StrEnum and datetime.UTC.
Reproducible backend test run from Docker:
make test-dockercd apps/analyst-console
npm install
npm run build
npm test- The worker runtime runs Bytewax inside the service process for local simplicity; distributed deployment tuning is still out of scope.
- The current MLflow service uses a local SQLite-backed metadata store suitable for demos, not HA production.
- Feedback-driven retraining is implemented as a controlled and manual workflow, not automatic promotion.
- Local Grafana notifications terminate at the API webhook sink rather than a real on-call channel.
- Kafka lag monitoring is not fully instrumented yet.
- The
Modelspage remains snapshot-style; the live polling emphasis is onOverviewandCases. - The local host shell often defaults to Python 3.10, so backend commands should be run through Docker or Python 3.11.
- Upgrade live polling to SSE or websockets when the product surface needs finer-grained updates
- Add richer backlog triage workflows and analyst assignment states
- Expand Kafka lag and consumer-group observability
- Add a replay mode for controlled historical stream demonstrations
- Add real screenshot assets and an optional short demo GIF
This repository is released under the MIT License.




