Candidate bench
Repo: https://github.com/lechmazur/writing
Owner: lechmazur (GitHub / @lechmazur on X)
Data status: ⚠️ needs-data — results live in README tables only; no standalone JSON or CSV file exists in the repository.
What it measures
Models are given a prompt containing 10 mandatory story elements (characters, objects, core concepts, attributes, actions, methods, settings, timeframes, motivations, and tone) and must weave all 10 into a short creative story. Multiple LLM judges score each story on element incorporation and overall quality. The primary metric is a comparison score reflecting how naturally and skillfully the model integrated all required elements.
Distinct from generic "write a story" prompts — the mandatory-element constraint tests structured creativity under constraints rather than free-form generation.
Why it fits BenchDirectory
- Personal/indie benchmark by a single GitHub user (same creator as the already-tracked lechmazur Sycophancy, Elimination Game, and Divergent Thinking candidates)
- Novel, opinionated question about a specific creative skill
- Very actively maintained: last updated July 14, 2026, adding GPT-5.6 Sol, Muse Spark 1.1, Grok 4.5
- Current frontier coverage: Claude Fable 5 (top), GPT-5.5/5.6, Gemini 3.5 Flash, DeepSeek V4 Pro, Kimi K2.6, Grok 4.3-4.5, Qwen 3.6-3.7
Data-reality check — NOT INGESTIBLE (yet)
Checked the repo on 2026-07-20: the root contains only images/, prompts_wc/, stories_wc/ directories and README.md. Leaderboard scores are rendered as tables inside the README — no machine-readable export exists. The published bundle contains prompt and story text files, not scores.
What the owner would need to publish
A CSV or JSON file at the repo root (e.g. results.csv) with at least: model, score (the comparison score), and ideally rank. Something like:
model,score,rank
claude-fable-5-high,3.3,1
gpt-5-6-sol-high,3.2,2
...
Once that file is published and accessible via raw GitHub URL, this bench is a strong build candidate.
Candidate bench
Repo: https://github.com/lechmazur/writing⚠️ needs-data — results live in README tables only; no standalone JSON or CSV file exists in the repository.
Owner: lechmazur (GitHub / @lechmazur on X)
Data status:
What it measures
Models are given a prompt containing 10 mandatory story elements (characters, objects, core concepts, attributes, actions, methods, settings, timeframes, motivations, and tone) and must weave all 10 into a short creative story. Multiple LLM judges score each story on element incorporation and overall quality. The primary metric is a comparison score reflecting how naturally and skillfully the model integrated all required elements.
Distinct from generic "write a story" prompts — the mandatory-element constraint tests structured creativity under constraints rather than free-form generation.
Why it fits BenchDirectory
Data-reality check — NOT INGESTIBLE (yet)
Checked the repo on 2026-07-20: the root contains only
images/,prompts_wc/,stories_wc/directories andREADME.md. Leaderboard scores are rendered as tables inside the README — no machine-readable export exists. The published bundle contains prompt and story text files, not scores.What the owner would need to publish
A CSV or JSON file at the repo root (e.g.
results.csv) with at least:model,score(the comparison score), and ideallyrank. Something like:Once that file is published and accessible via raw GitHub URL, this bench is a strong build candidate.