# Install & setup
uv sync
export OPENROUTER_API_KEY="your-key"
# Run games
uv run based codenames run --red gpt4 --blue claude
uv run based chainlex run --model-away gpt-4o --model-home gemini-3-flash
uv run based connections run --model gemini-flash --puzzles 10
# Tournament evaluation
uv run based chainlex eval --all --dry-run # Preview schedule
uv run based chainlex eval --all # Full round-robin (16 threads)
uv run based chainlex eval --add-model new-model # Add model to existing results
# Analytics / Leaderboard
uv run based analytics leaderboard -r logs/chainlex/eval/detailed_results.csv
# DSPy optimization
uv run based chainlex optimize --model gemini-3-flash --num-train 50
uv run based chainlex optimize --model gemini-3-flash --blend # Round-robin across 3 models
uv run based chainlex optimize --model gemini-3-flash --budget small # light/medium/large/insane
uv run based chainlex deploy-prompts
# Development
uv run pytest
uv run black . && uv run isort .- Board Setup: 25 words (9 red, 8 blue, 7 bystanders, 1 assassin)
- Turn Loop: Teams alternate Spymaster → Operative phases
- Win: Find all agents OR opponent hits assassin
- Board Setup: 16 words (8 friendly, 7 bystanders, 1 assassin)
- Away Turn: First player gives clue + guesses (blind to opponent)
- Home Turn: Second player knows opponent's score, adapts strategy
- Scoring: Triangular (1+2+3+...), bystander=-1, assassin=instant loss
- Win: Higher score wins; both hit assassin = tie
All scoring and parsing logic lives in chainlex/game_engine.py:
GameEngine.score_guesses()- Canonical scoring functionGameEngine.parse_clue_from_response()- Canonical clue parserGameEngine.parse_guesses_from_response()- Canonical guess parserBoardStatedataclass for consistent board representation
This ensures optimizer and game use identical rules (no drift).
Each matchup plays 4 games on 2 puzzles (balanced difficulty):
- Hard puzzle × 2 games (home/away swap)
- Easy puzzle × 2 games (home/away swap)
This ensures models are tested on both difficulty levels for fair comparison.
- Stateless AI Calls: Each OpenRouter request independent
- External Prompts: All prompts in Markdown files (
prompts/) - Home/Away Mechanic: Second player advantage amplifies model intelligence differences
- Append Mode:
--add-modelappends to existing results.csv - Single Source of Truth: All game logic (scoring, parsing) in
game_engine.py - No Optimizer Drift: Optimizer uses GameEngine, not duplicate logic
- Puzzle Pool Separation: Training puzzles never leak to eval/run
- Difficulty Balance: Eval uses 1 hard + 1 easy puzzle per matchup
- Type hints throughout
richfor console outputtyperfor CLI- JSONL for structured logs
uv run based codenames prompt spymaster --seed 42 --team red
uv run based chainlex prompt clue_giver --seed 42- Edit
shared/inputs/model_mappings.yml - Test:
uv run based codenames run --red NEW_MODEL --blue claude
- Edit
prompts/*.mdfiles - Template vars:
{{BOARD}},{{CLUE}},{{HEAD_TO_HEAD_CONTEXT}},{{AVAILABLE_WORDS}} - Test with
--verboseflag - Critical: Guesser prompt MUST include template section at end with
{{AVAILABLE_WORDS}},{{CLUE}},{{NUMBER}} - Parser looks for
## Guessessection - ensure prompts instruct model to use this format - Parser strips
*, backticks, and markdown formatting from extracted words
--verbosefor full AI exchanges- Check
logs/directory - Use
--seedfor reproducibility - Query controllog events for full request/response text:
SELECT payload_json.request_text, payload_json.response_text FROM controllog.events WHERE kind = 'model_prompt' OR kind = 'model_completion'
| Term | Codenames | ChainLex-1 |
|---|---|---|
| Clue giver | Spymaster | Clue Giver |
| Guesser | Operative | Guesser |
| Target word | Agent | Friendly Word |
| Neutral | Bystander (-0) | Bystander (-1) |
| Instant loss | Assassin | Assassin (instant loss) |
| Advantage | First turn | Home (2nd, knows score) |
codenames/ # Codenames game
chainlex/ # ChainLex-1 game
├── game.py # Full game orchestration
├── game_engine.py # Shared scoring/parsing (SINGLE SOURCE OF TRUTH)
├── player.py # AI player wrapper (uses GameEngine)
├── puzzle_generator.py # Semantic clustering puzzle generation
├── puzzle_loader.py # Training/eval puzzle pool loader
├── optimization/ # DSPy prompt optimization (uses GameEngine)
│ ├── optimize.py # GEPA/MIPROv2 optimization
│ ├── modules.py # DSPy modules (uses GameEngine parsing)
│ └── metrics.py # Scoring metrics (uses GameEngine scoring)
├── prompts/ # Role prompts (clue_giver.md, guesser.md)
├── inputs/ # Puzzle pools
│ ├── puzzles_training.yaml # 50 training puzzles (optimizer only)
│ ├── puzzles_eval.yaml # 50 eval puzzles (eval/run commands)
│ └── word_pool.yaml # Words with semantic categories
└── optimized_prompts/ # DSPy output
connections/ # Connections game
shared/ # Infrastructure (controllog, adapters, utils)
logs/
├── controllog/ # Unified analytics (ALL games write here)
│ └── YYYY-MM-DD/ # Date-partitioned events.jsonl + postings.jsonl
├── chainlex/ # ChainLex game-specific logs (box_scores, etc.)
├── codenames/ # Codenames game-specific logs
└── eval/ # Tournament results
ChainLex uses separate puzzle pools to prevent data leakage during optimization:
| Pool | File | Size | Used By |
|---|---|---|---|
| Training | puzzles_training.yaml |
50 | optimize command only |
| Eval | puzzles_eval.yaml |
50 | eval, run, cost-estimate |
Puzzle Selection (eval command):
- Selects 1 hard + 1 easy puzzle per matchup
- Each puzzle played twice (home/away swap) = 4 games total
- Ensures balanced difficulty testing
Puzzle Generation (generate-puzzles command):
- Uses semantic embeddings (sentence-transformers) for word clustering
- Validates puzzles for cohesion, bystander confusion, assassin proximity
- Difficulty tiers: easy (0.45-0.6 cohesion), medium (0.35-0.45), hard (0.25-0.35)
# Generate new puzzle pools
uv run based chainlex generate-puzzles
# List puzzles in a pool
uv run based chainlex list-puzzles --pool training
# List only hard puzzles
uv run based chainlex list-puzzles --pool eval --difficulty hardIMPORTANT: All games must initialize controllog with Path("logs") (the top-level logs directory):
cl.init(project_id="chainlex", log_dir=Path("logs")) # ✅ Correct
cl.init(project_id="chainlex", log_dir=log_dir) # ❌ Wrong if log_dir is game-specificThe SDK automatically creates logs/controllog/<date>/ subdirectories. This ensures uv run based analytics upload finds all game data without needing --log-path flags.
Game-specific files (box_scores, game_metadata, play_by_play) can still go to game-specific directories.
| Event | Purpose | Key Fields |
|---|---|---|
model_prompt |
AI request | request_text, prompt_tokens, model |
model_completion |
AI response | response_text, completion_tokens, wall_ms, cost_money |
state_move |
State transition | from_, to (NEW→WIP→DONE) |
game_complete |
Game summary | outcome, winner_model, scores, wall_ms |
Emitted at game end for leaderboards and analytics:
cl.game_complete(
task_id="game:abc123",
game_id="abc123",
model_away="claude-3",
model_home="gpt-4",
outcome="model_away", # or "model_home" or "tie"
winner_model="claude-3",
score_away=15,
score_home=10,
margin=5,
correct_guesses_away=4,
correct_guesses_home=3,
total_guesses=8,
wall_ms=45000,
cost_money=0.01,
)This enables:
- Leaderboards: Query by outcome/winner_model
- Head-to-head matrices: Group by (model_away, model_home) pairs
- Efficiency metrics: cost_money/correct_guesses, wall_ms/total_guesses
TODO: Add game_complete events to Codenames (#4) and Connections (#5)
The analytics leaderboard command generates Bradley-Terry ratings using arena-rank.
| File | Columns | Features |
|---|---|---|
results.csv |
model_a, model_b, winner |
Basic BT ratings |
detailed_results.csv |
model_home, model_away, winner, ... |
BT ratings + home/away splits |
Ties are handled per arena-rank defaults:
winner="tie"→ outcome 0.5 (half-win for each model)winner="model_a"→ outcome 1.0 (model_a wins)winner="model_b"→ outcome 0.0 (model_b wins)
Win% in tables uses (wins + 0.5*ties) / total_games to match BT model.
# In shared/cli_analytics.py
BradleyTerry(n_competitors=len(dataset.competitors), init_rating=1600)init_rating=1600: Baseline rating (matches standard ELO)- Variance will be high with few games per matchup (aim for 20+ per pair)
leaderboard.csv: Full rankings with CIleaderboard.png: Forest plot (dot + CI range, color-coded by tier)
# Run all tests
uv run pytest tests/
# Run specific test files
uv run pytest tests/test_chainlex_game.py -v # ChainLex game logic
uv run pytest tests/test_chainlex_player.py -v # ChainLex player/parsing
uv run pytest tests/test_controllog.py -v # Controllog SDK
uv run pytest tests/test_game.py -v # Codenames game logic| File | Tests | Coverage |
|---|---|---|
test_chainlex_game.py |
24 | Board setup, scoring, clue validation, controllog |
test_chainlex_player.py |
22 | Response parsing, metadata storage, board formatting |
test_controllog.py |
9 | game_complete, model_prompt/completion text |
test_game.py |
7 | Codenames game logic |
test_metadata_logging.py |
6 | Adapter cost extraction |
OPENROUTER_API_KEY: Required for AI modelsMOTHERDUCK_DB: Optional, for analytics upload