Four AI approaches to predicting Fantasy Premier League player points — from traditional ML to fine-tuned LLMs, automated research, and Claude Haiku.
The Experiment · How It Works · Auto-Research · Claude Haiku · Live Results · Project Structure · Progress · Reproduce
Four AI approaches. Same prediction problem. Each with different trade-offs — and the winner came from an automated research loop that designed its own architecture.
Approach 1 — XGBoost (Traditional ML): Feed structured player stats into a gradient-boosted decision tree. Fast, proven, and the industry standard for tabular prediction tasks.
Approach 2 — Auto-Research XGBoost (AI-Designed): An automated experimentation loop where Claude Code acts as the researcher — proposing changes, testing them, and only keeping improvements. Discovered a two-stage architecture (classifier + regressor) that beat the hand-tuned original. 76 experiments, 12 committed improvements.
Approach 3 — Fine-Tuned LLM (Llama 3.2 3B): Take an open-source language model, fine-tune it with LoRA on natural language descriptions of player stats, and ask it to predict points. Runs entirely on a Mac Mini M4 using Apple's MLX framework — no cloud GPU needed.
Approach 4 — Claude Haiku (API-Based LLM): Send player stats and other models' predictions to Claude Haiku via the Anthropic API. No fine-tuning — just prompt engineering with structured context. Provides an independent third-party signal plus natural language reasoning for each prediction.
The question isn't just "which is more accurate?" It's: when does each approach add value, and what are the real-world trade-offs?
All data comes from the free, public FPL API. The pipeline fetches player stats, fixture data, and gameweek histories for 800+ players, then engineers 20 features per player-gameweek:
| Feature Type | Examples |
|---|---|
| Rolling form | Points avg (3/5 GW), minutes avg, bonus avg, ICT index avg |
| Season stats | Goals per 90, assists per 90, clean sheet % |
| Fixture context | Home/away, fixture difficulty rating (1-5) |
| Team context | Team goals scored/conceded (last 3 GW), opponent goals conceded (season avg) |
All rolling features are computed using only prior gameweeks — no future data leakage.
┌─────────────────────────────┐
│ 8,468 player-gameweek │
│ training examples │
└──────────┬──────────────────┘
│
┌──────────────┬───────────┴──────────┬──────────────┐
▼ ▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ XGBoost │ │ Auto-Research│ │ Llama 3.2 3B │ │ Claude Haiku │
│ (original) │ │ XGBoost │ │ + LoRA v3 │ │ (API) │
│ │ │ │ │ │ │ │
│ Single │ │ Classifier + │ │ Chat prompts │ │ Structured │
│ regressor │ │ Regressor │ │ → Int pred │ │ prompt + │
│ │ │ (2-stage) │ │ → Smoothed │ │ reasoning │
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │ │
└────────┬───────┴───────────┬───────┘ │
│ └────────────────┬───────┘
▼ ▼
┌─────────────────────┐ ┌──────────────────┐
│ Evaluation Harness │ │ Combined Value │
│ MAE, RMSE, ±1/±3 │ │ Score (avg of │
│ Category breakdown │ │ all 4 models) │
└─────────────────────┘ └──────────────────┘
Not just "what's the MAE?" — a structured eval suite with named test scenarios:
| Category | What it tests |
|---|---|
| Fixture context | Does the model understand home vs away, easy vs hard fixtures? |
| Positional | Can it handle GK/DEF/MID/FWD differently? |
| Edge cases | Low-minutes players, hot streaks, weak teams |
| High value | Captain picks, differentials — the predictions that matter |
| Human confidence | Accuracy on cases a human labelled as easy vs hard to predict |
Every experiment run is compared against a saved baseline, with automatic regression detection.
Evaluated on 243 player-gameweeks from GW31+ (held-out test set):
| Model | MAE | RMSE | Within ±1 pt | Within ±3 pts | Inference Time |
|---|---|---|---|---|---|
| XGBoost (auto-research) | 1.90 | 3.00 | 49.8% | 79.8% | <1s |
| XGBoost (original) | 2.11 | 2.81 | 26.7% | 82.3% | <1s |
| Few-Shot LLM (3 examples) | 2.16 | 3.32 | 55.6% | 76.1% | 370s |
| Chain-of-Thought LLM | 2.20 | 3.60 | 60.1% | 73.7% | 845s |
| Fine-Tuned LLM v3 | 3.35 | 4.46 | 39.1% | 65.8% | 148s |
| Claude Haiku (API) | — | — | — | — | ~60s |
| Zero-Shot LLM (baseline) | 11.54 | 14.22 | 2.1% | 10.3% | 133s |
Note: Claude Haiku eval metrics are not directly comparable as it receives other models' predictions as context. Its value is as an independent signal with natural language reasoning, not as a standalone predictor.
XGBoost by position:
| Position | MAE | Samples |
|---|---|---|
| GK | 1.72 | 16 |
| DEF | 2.39 | 76 |
| MID | 1.83 | 118 |
| FWD | 2.64 | 33 |
MAE comparison (lower is better)
─────────────────────────────────────────────────────────
Auto-Research ████████████████████ 1.90 <-- winner (AI-designed)
XGBoost (orig) ██████████████████████ 2.11
Few-Shot LLM ██████████████████████▌ 2.16
Chain-of-Thought ███████████████████████ 2.20
Claude Haiku (uses other models as context — not standalone)
Fine-Tuned v3 ███████████████████████████████████▌ 3.35
Zero-Shot ████████████████████████████████████████████████████████████ 11.54
0 2 4 6 8 10 12
The most interesting story in this project isn't the final numbers — it's how we got there. Fine-tuning the LLM took 3 iterations to fix a critical training data bug, revealing that data quality matters more than model capacity.
FPL points follow a heavily skewed distribution. Most players score 1-2 points in any given gameweek:
Training data distribution (8,225 player-gameweeks)
───────────────────────────────────────────────────
0 pts ████████████████ 12%
1 pt ████████████████████████████ 22%
2 pts ████████████████████████████████████████████████ 36% <-- mode
3 pts ████████████ 9%
4 pts ████████████ 9%
5 pts ██████ 5%
6 pts ████ 3%
7 pts ██ 2%
8+ pts ██ 2%
This is real-world class imbalance: 58% of training examples are 1-2 points, while high-value predictions (5+ pts) make up just 12%.
Config: 600 iterations, 8 LoRA layers, learning rate 1e-5, rank 8
What happened: The model learned a shortcut. Instead of understanding player form, fixture difficulty, and positional context, it simply learned to always predict 2 — the most common value in the training data.
v1 prediction distribution (243 test players)
──────────────────────────────────────────────
0 ░
1 ░
2 █████████████████████████████████████████████████████████████ ~95%
3 ░
4 ░
5+ ░
"Every player gets 2 points"
Why it happened: With 58% of training data concentrated at 1-2 points, predicting "2" every time minimises the loss function. The model found a local minimum that's hard to escape with standard training. This is a well-known problem in ML called mode collapse — the model collapses to the most frequent class.
Lesson: A low training loss doesn't mean the model learned anything useful. Always inspect the prediction distribution, not just the aggregate error.
Config: 400 iterations, 8 LoRA layers, learning rate 2e-5 (doubled)
Key change: Added balance_training_data() to the prompt generation pipeline. This function oversampled underrepresented point ranges and capped overrepresented ones:
Before (raw): 1-2 pts = 58% → After (balanced): 1-2 pts = 37%
3-4 pts = 18% 3-4 pts = 25%
5+ pts = 12% 5+ pts = 38%
Result: The model broke free from mode collapse and started differentiating:
v2 prediction distribution
──────────────────────────
0 ██████████ ~15% (low-form players)
1 ████████████████ ~25%
2 ████████████████████ ~30%
4 ████████████████ ~25% (high-form players)
5+ ░░ (still missing!)
"Differentiates, but only uses 4 values"
What improved: Low-form players in tough fixtures got 0-1, high-form players got 4. Real differentiation.
What was still broken: The model only predicted {0, 1, 2, 4} — no predictions above 4. Premium players having big weeks (8-12 pts) were invisible.
Lesson: Coarse bucket balancing (grouping 5+ points together) taught the model that "high" exists, but not how high. The resolution of the balancing matters.
Config: 500 iterations, 16 LoRA layers (doubled), learning rate 3e-5
Two changes made together:
-
Per-point balancing: Instead of grouping into coarse buckets, created exactly 333 examples for each point value 0-11. The model saw equal representation of every outcome.
-
Doubled LoRA capacity: 8 → 16 layers. More learnable parameters to handle the finer-grained prediction task.
v3 prediction distribution
──────────────────────────
0 ████████ ~12% (bench warmers, injured)
1 ████████████ ~18% (low form, tough fixtures)
2 ████████████████ ~25% (average players)
4 ████████████████ ~22% (good form)
9 ██████ ~10% (premium players, easy fixtures)
10 ██████ ~8% (captaincy picks)
"Full 0-10 range with meaningful differentiation"
Trade-off: Raw MAE went from ~2.5 (v2) to 3.35 (v3). The wider prediction range means bigger misses when wrong, but the model now ranks players meaningfully — essential for FPL where you need to pick the best 11, not just predict the average.
Iteration LoRA LR Iters Data Balance Unique MAE
Layers Values
─────────────────────────────────────────────────────────────────────────
v1 8 1e-5 600 None (raw) 1 (!) ~2.5*
v2 8 2e-5 400 Bucket-balanced 4 ~2.8
v3 16 3e-5 500 Per-point (333 ea) 6 3.35
─────────────────────────────────────────────────────────────────────────
* v1 MAE looks decent because predicting "2" for everyone
is close to the mean — but the model learned nothing
Model capacity vs data quality impact
──────────────────────────────────────
v1 ──[fix data]──> v2 ──[fix data + add capacity]──> v3
Data quality fix (v1→v2): Unlocked differentiation (1 → 4 unique values)
Capacity + data fix (v2→v3): Unlocked full range (4 → 6 unique values)
Takeaway: Data quality was the bottleneck, not model size.
Doubling LoRA layers without fixing data (v1) would have
just produced a more confident "always predict 2" model.
The v3 fine-tuned model predicts discrete integers {0, 1, 2, 4, 9, 10}. For production use, we blend the LLM signal with player features to produce realistic decimal predictions:
Raw LLM output Smoothed output
──────────────── ─────────────────
6 unique integers → 66 unique decimal values
Range: 0-10 → Range: 0.5-9.5
Mean: ~2.8 → Mean: ~3.05
Formula:
┌─────────────────────────────────────────────────────────┐
│ 1. Dampen LLM: llm * 0.6 + 1.0 (clip to 0.5-7.5) │
│ 2. Feature est: form*0.5 + form5*0.3 + ict*0.15 │
│ + fixture_adj + attack + cs_bonus │
│ 3. Blend: 40% LLM + 30% form + 30% features │
│ 4. Clip & round: 0.5-10.0, 1 decimal place │
└─────────────────────────────────────────────────────────┘
Why blend? The LLM provides directional signal (who will score high vs low), while form stats provide calibration (realistic point ranges). The blend captures both.
The most unexpected result: giving the base model (no fine-tuning) just 3 examples in the prompt achieved MAE 2.16 — within 0.05 of XGBoost's 2.11.
Accuracy vs effort trade-off
────────────────────────────
MAE
12 │ x Zero-Shot (11.54)
│ "The model knows nothing about FPL"
│
4 │
│ x Fine-Tuned v3 (3.35)
3 │ "Wider range, bigger misses"
│
│ x CoT (2.20)
2 │ x Few-Shot (2.16) x XGBoost (2.11)
│ "3 examples in prompt" "500 trees, 8K rows"
│
0 └───┬──────┬──────┬──────┬──────┬──────────────
0 1h 2h 3h 4h Dev time
(prompt (fine- (hyper- (full
eng) tune) tune) pipeline)
Implication for the "build vs prompt" decision:
- If you need quick, good-enough predictions: few-shot prompting gets you 97% of XGBoost's accuracy with zero training.
- If you need best possible accuracy: Auto-Research XGBoost wins on MAE (1.90) and is 370x faster than any LLM at inference.
- If you need player ranking/differentiation: fine-tuned LLM provides the widest prediction range.
- If you need explainability: Claude Haiku provides natural language reasoning for each prediction.
- If you want the most robust signal: combine all four — the dashboard averages their value scores, and model agreement highlights high-conviction picks.
Inspired by Andrej Karpathy's "vibe coding" auto-research concept, we built an automated experimentation loop where Claude Code acts as the researcher — proposing changes to the model, testing them, and only keeping improvements.
Instead of manually tweaking hyperparameters, the system runs a tight loop:
┌────────────────────────────────────────────────────────┐
│ Claude Code reads current model code & baseline MAE │
│ │ │
│ ▼ │
│ Proposes a change to train.py │
│ (architecture, hyperparams, blending) │
│ │ │
│ ▼ │
│ Runs autoresearch/run_loop.py │
│ (train → evaluate → compare) │
│ │ │
│ ┌────────┴────────┐ │
│ ▼ ▼ │
│ Improved? No better? │
│ Git commit Discard & │
│ + update try next │
│ baseline idea │
│ │ │ │
│ └────────┬────────┘ │
│ ▼ │
│ Repeat │
└────────────────────────────────────────────────────────┘
The auto-research loop found a two-stage architecture that outperforms the original single-model approach:
- Classifier — "Will this player score any points?" (handles benched/injured players)
- Regressor — "How many points?" (predicts the actual total)
- Combine —
probability_of_playing^1.2 × predicted_points
It also discovered that switching from squared error to absolute error, using shallower trees, and applying probability sharpening all improved accuracy.
| Original XGBoost | Auto-Research XGBoost | |
|---|---|---|
| Architecture | Single regressor | Classifier + Regressor (two-stage) |
| Objective | reg:squarederror |
reg:absoluteerror |
| Trees | 500 | 1,500 (reg) + 500 (clf) |
| Tree depth | 6 | 4 (reg) / 2 (clf) |
| Prediction formula | predict(X) |
play_prob^1.2 × predict(X) |
| MAE | 2.22 | 1.90 (14.6% better) |
| Experiments run | — | 76 (12 improvements committed) |
The full experiment trail is preserved in git history — every commit tagged [autoresearch] shows exactly which change caused which improvement.
Three small modules in autoresearch/:
| File | Purpose |
|---|---|
prepare.py |
Loads data, splits train/test, computes metrics |
train.py |
Model architecture & hyperparameters (the thing being tweaked) |
run_loop.py |
Orchestrator: train → evaluate → compare → exit code 0/1 |
The fourth prediction model uses Claude Haiku via the Anthropic API. Unlike the other models, Claude receives other models' predictions as context alongside the player stats — making it a meta-predictor that can synthesise signals from XGBoost and the LLM.
- No training required — pure prompt engineering with structured player data
- Natural language reasoning — each prediction comes with an explanation (e.g. "High form at home against a leaky defence, but rotation risk due to midweek fixture")
- Independent signal — trained on different data from the other models, so disagreements are informative
- Batch or individual — can predict 15 players at once or one at a time
Claude predictions are generated separately (python models/predict_claude.py) and merged into the main predictions CSV. The dashboard shows Claude alongside the other three models, and the combined value score averages all available models.
We're tracking how each model performs on real, unseen gameweeks — predictions made before the deadline, scored after the matches. Full analysis, per-player breakdowns, and cumulative standings are in RESULTS.md.
Cumulative standings after GW32:
| Rank | Model | Total Points | Avg MAE |
|---|---|---|---|
| 1 | Claude Haiku | 53 | 4.20 |
| 2 | LLM (Llama 3.2 3B) | 44 | 4.81 |
| 3 | Auto-Research XGBoost | 42 | 3.78 |
| 4 | XGBoost (Original) | 18 | 1.95 |
Early days — one gameweek is noise. See RESULTS.md for the full story, including why XGBoost's low MAE is misleading and how captain selection swung the results by 8+ points.
- Phase 1: Data Pipeline — FPL API fetcher, feature engineering (20 features), prompt generation
- Phase 2: XGBoost Baseline — Trained and evaluated (MAE 2.11)
- Phase 3: LLM Fine-Tuning — 3 iterations of LoRA fine-tuning on Llama 3.2 3B with MLX (v1: mode collapse, v2: partial range, v3: full 0-10 range)
- Phase 4: Eval Framework — Category-based eval suite (14 named cases across 4 categories), output quality checks, regression detection
- Phase 5: Comparison Dashboard — Multi-page Streamlit app with predictions table, fixture lookahead, squad builder, and transfer advisor
- Phase 6: Write-Up — Product assessment, "would I ship this?" brief, build-vs-buy-vs-prompt framework
- Phase 7: Notebooks & Docs — 5 documented Jupyter notebooks covering data exploration, training, fine-tuning, evals, and final comparison
- Phase 8: Auto-Research — Automated experimentation loop using Claude Code. Two-stage classifier+regressor architecture, 76 experiments, MAE improved from 2.22 to 1.90 (14.6% gain)
- Phase 9: Claude Haiku — API-based predictions with natural language reasoning. Independent signal that enriches the combined value score across all 4 models
data/
fetch_fpl.py Fetches all data from FPL API
build_features.py Engineers 20 features per player-gameweek
build_prompts.py Generates LLM training prompts (JSONL + MLX chat format)
archive_predictions.py Archives current GW predictions before generating next GW
predictions/ Archived per-GW prediction CSVs (gw32_predictions.csv, ...)
processed/ Feature CSV + prediction outputs
mlx/ MLX-formatted train/valid/test splits
raw/ Raw JSON from FPL API (players, fixtures, teams, gameweeks)
models/
train_xgboost.py XGBoost training script
predict_next_gw.py Next-GW prediction pipeline (XGBoost + Auto-Research + LLM)
predict_claude.py Claude Haiku predictions via Anthropic API
predict_llm.py LLM evaluation harness (4 strategies)
xgboost_fpl.json Pre-trained XGBoost model
llama-3.2-3b/ Base model (4-bit quantised)
fpl-lora-adapter-v3/ Fine-tuned LoRA weights (v3 — per-point balanced)
autoresearch/
prepare.py Data loading, train/test split, metric evaluation
train.py Model architecture & hyperparameters (classifier + regressor)
run_loop.py Experiment orchestrator: train → eval → compare → report
baseline_mae.json Current best performance metrics
eval/
eval_suite.yaml Named eval scenarios (4 categories, 14 cases)
build_cases.py Generates eval case files from YAML + test data
run_eval.py Single-command eval runner with category scoring
compare.py Regression detection between eval runs
checks/output_checks.py LLM output quality validators
baseline/scores.json Current best results to compare against
results/ Timestamped eval run outputs
notebooks/
01_data_exploration.ipynb Dataset analysis: distributions, features, correlations
02_xgboost_training.ipynb XGBoost training, feature importance, residual analysis
03_llm_finetuning.ipynb LoRA fine-tuning journey: 3 iterations, mode collapse fix
04_eval_framework.ipynb Category-based eval suite, regression detection
05_comparison_and_assessment.ipynb Final head-to-head comparison and product assessment
writeup/
assessment.md Full findings and product leader takeaways
product_brief.md "Would I ship this?" decision framework
ui/
app.py Entry point — multi-page navigation shell
data_loader.py Cached data loading + fixture lookahead
styles.py FPL-themed CSS styling
components.py Reusable HTML components (badges, cards, etc.)
pages/
1_Predictions.py Main predictions table with filters + next 2 fixtures
2_My_Team.py Squad builder + multi-transfer advisor
3_Best_XI.py Optimal predicted XI with pitch view
4_Results.py Gameweek results — predicted vs actual comparison
results_logic.py Scoring logic for model evaluation
docs/
setup-explained.html Plain-English guide to everything we built
git clone https://github.com/tom-barkan/FPL-MLprediction-model.git
cd FPL-MLprediction-model
# Set up environment
brew install python@3.11 libomp
python3.11 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# Fetch data from FPL API (~15 min due to rate limiting)
python data/fetch_fpl.py
# Build features and prompts
python data/build_features.py
python data/build_prompts.py
# Train XGBoost
python models/train_xgboost.py
# Fine-tune LLM (requires Apple Silicon)
pip install mlx mlx-lm
huggingface-cli download mlx-community/Llama-3.2-3B-Instruct-4bit --local-dir models/llama-3.2-3b
mlx_lm.lora --model models/llama-3.2-3b --data data/mlx --train --iters 600 --adapter-path models/fpl-lora-adapter-v2
# Generate next-GW predictions (XGBoost + Auto-Research + LLM)
python models/predict_next_gw.py
# Add Claude Haiku predictions (requires ANTHROPIC_API_KEY)
python models/predict_claude.py
# Launch the dashboard
streamlit run ui/app.pyThe Streamlit dashboard has four pages:
Predictions — Sortable/filterable table showing all 4 models' predictions for 344 players. Model filter toggle, next 2 fixtures with difficulty ratings, player deep dive with per-model cards, and model agreement/disagreement analysis.
My Team — Build your 15-player FPL squad and get transfer recommendations. Squad validation, auto-picked starting XI, and a multi-transfer planner with four strategy tabs (Safe, Differential, Form, Fixture).
Best XI — Optimal predicted starting XI within budget constraints, shown on a football pitch with captain/vice-captain badges and fixture difficulty indicators.
Results — Gameweek-by-gameweek comparison of each model's Best XI against actual FPL points. Pitch views with actual scores, metrics comparison table, and cumulative model performance tracking. See RESULTS.md for the full write-up.
| Component | Spec |
|---|---|
| Machine | Mac Mini M4 |
| RAM | 16GB unified memory |
| Cloud GPU | Not required |
| Fine-tuning time | ~20 minutes |
| Inference (full test set) | ~5 minutes |
| Tool | Purpose |
|---|---|
| MLX + mlx-lm | Local LLM fine-tuning and inference on Apple Silicon |
| Llama 3.2 3B | Base model for fine-tuning (4-bit quantised, ~2GB) |
| XGBoost | Traditional ML baseline + auto-research optimised two-stage model |
| Claude Haiku | API-based predictions with natural language reasoning |
| Claude Code | Automated research loop — proposes, tests, and commits model improvements |
| FPL API | Free, public source for all player and fixture data |
| Streamlit | Comparison dashboard |
| Jupyter | Documented experiment notebooks |