End-to-end DLOps showcase: a Vision Transformer fine-tuned to classify 30-day candlestick charts into 3 forward-direction classes (up / sideways / down), wrapped in a reproducible DVC pipeline, MLflow tracking, FastAPI inference, Gradio demo, Prometheus/Grafana observability, Docker + Kubernetes deploy, and full GitHub Actions CI/CD.
Try it on Hugging Face Spaces →
Upload a candlestick chart → get a directional prediction with class probabilities and an LLM-generated technical-analysis explanation. Runs on CPU, free-tier hosted, no signup required.
The deployed Space uses a simplified single-file app.py that loads the trained checkpoint directly (see the Space repo). The full FastAPI + Gradio + observability architecture described below is the local / production stack.
Given a 30-day candlestick chart of an S&P 500 ticker, predict the direction of the 5-day forward return (±2% threshold → up / sideways / down). An OpenAI-backed explainer pairs each prediction with a one-paragraph rationale in plain English.
Three representative chart windows — one from each class — show what the ViT actually classifies.
| Up | Sideways | Down |
|---|---|---|
| ascending staircase | flat / choppy | descending cliff |
The system has four tiers: user-facing UI (Gradio), inference (FastAPI with ViT vision model + OpenAI explainer), training (DVC + MLflow), and infrastructure (GitHub Actions CI/CD, Terraform-provisioned GKE).
- DVC pipeline with parameter + dependency hashing for reproducible reruns
- MLflow for hyperparameter, metric, and artifact tracking
- Time-aware train/val/test split — chronological, no random shuffling, no leakage
- Class weighting in the loss function to counter label imbalance
- FastAPI inference server with
/healthz,/predict,/metrics - Prometheus + Grafana observability stack via docker-compose
- Gradio demo UI sitting alongside the API
- GitHub Actions for CI (lint, type-check, test) and CD (image build, GKE rollout)
- Terraform for the GCP foundation: GKE Autopilot, GCS DVC remote, Artifact Registry
- Kubernetes manifests with HPA, Ingress, ConfigMap, Secret template
- Docker multi-stage builds with GitHub Actions cache for fast CI rebuilds
- Memory-mapped checkpoint loading for robust serving on fragmented systems
The DVC DAG defines six reproducible stages from raw CSV to trained model:
load_ohlcv → label_windows → render_charts → build_dataset → train → evaluate
Single-run honest reporting from the held-out chronological test set.
| Class | Precision | Recall | F1 |
|---|---|---|---|
| up | 0.40 | 0.85 | 0.54 |
| sideways | 0.41 | 0.20 | 0.27 |
| down | 0.00 | 0.00 | 0.00 |
| Macro | – | – | 0.27 |
Test accuracy: 0.40 (vs. 0.33 uniform-random baseline and 0.38 majority-class baseline, i.e. always predicting "up")
This project is a DLOps demonstration, not an alpha-generating signal.
- Two CPU epochs is far from convergence — ViT-base needs many more epochs (or a GPU) to specialize from ImageNet pretraining onto chart patterns
- The down class collapsed at 0.0 recall — class weighting helped marginally; the real fix is freezing the backbone or migrating training to GPU
- The pipeline, reproducibility, and operational story are the actual portfolio value, not the headline accuracy
Next steps: train on GPU for more epochs, compare against a majority-class baseline and a non-image model (e.g. gradient boosting on raw OHLCV returns), and try volatility-adjusted labels instead of a fixed ±2% threshold.
Every training run is logged with full provenance: source file path, git commit hash, hyperparameters, metrics, and artifacts.
# 1. Install dev + runtime requirements
make dev
# 2. Set up environment (.env)
cp .env.example .env
# Edit .env to add OPENAI_API_KEY, HF_TOKEN, DVC_GDRIVE_FOLDER_ID
# 3. Drop the Kaggle dataset into data/raw/
# https://www.kaggle.com/datasets/andrewmvd/sp-500-stocks
# 4. Run the full pipeline (data → training → evaluate)
dvc repro
# 5. Serve the API
make serve
# 6. Launch the demo (in a separate terminal)
python -m src.ui.gradio_appThen open:
- API docs: http://localhost:8000/docs
- Gradio demo: http://localhost:7860
- MLflow UI:
mlflow ui→ http://localhost:5000
Or skip the local setup entirely and use the live demo on Hugging Face Spaces.
| Layer | Tool |
|---|---|
| Modeling | PyTorch + HuggingFace ViT |
| Experiment tracking | MLflow |
| Data versioning | DVC (Google Drive remote) |
| Serving | FastAPI + Uvicorn |
| Demo UI | Gradio |
| LLM explainer | OpenAI gpt-4o-mini |
| Observability | Prometheus + Grafana |
| Container | Docker + docker-compose |
| Orchestration | Kubernetes (GKE Autopilot) |
| Cloud | Google Cloud Platform |
| Infrastructure as Code | Terraform |
| CI/CD | GitHub Actions + Jenkins |
.
├── src/ # FastAPI, training, inference, UI, models, data
├── config/ # Runtime configs (model, serving, data)
├── data/ # DVC-tracked datasets (raw / interim / processed / splits)
├── docker/ # Dockerfiles + docker-compose stack + prometheus.yml
├── k8s/ # Kubernetes manifests for GKE
├── terraform/ # Cloud foundation (cluster, bucket, registry)
├── jenkins/ # On-prem alternative CI pipeline
├── scripts/ # Deploy and DVC setup scripts
├── tests/ # Pytest suite (data, model, API, pipeline smoke)
├── notebooks/ # Exploration
├── docs/ # Architecture diagrams, demo captures, screenshots
├── .github/workflows/ # CI / CD / DVC-repro automation
├── params.yaml # DVC-tracked hyperparameters
└── dvc.yaml # DVC pipeline definition
See LICENSE.
- Prakhar Srivastava
- Data Scientist, Business Analyst & AI Engineer | Machine Learning, Deep Learning & AI Automation Enthusiast
















