Skip to content

Candidate: LLM Creative Story-Writing Benchmark #8

Description

@Dan-Cleary

Candidate bench

Repo: https://github.com/lechmazur/writing
Owner: lechmazur (GitHub / @lechmazur on X)
Data status: ⚠️ needs-data — results live in README tables only; no standalone JSON or CSV file exists in the repository.

What it measures

Models are given a prompt containing 10 mandatory story elements (characters, objects, core concepts, attributes, actions, methods, settings, timeframes, motivations, and tone) and must weave all 10 into a short creative story. Multiple LLM judges score each story on element incorporation and overall quality. The primary metric is a comparison score reflecting how naturally and skillfully the model integrated all required elements.

Distinct from generic "write a story" prompts — the mandatory-element constraint tests structured creativity under constraints rather than free-form generation.

Why it fits BenchDirectory

  • Personal/indie benchmark by a single GitHub user (same creator as the already-tracked lechmazur Sycophancy, Elimination Game, and Divergent Thinking candidates)
  • Novel, opinionated question about a specific creative skill
  • Very actively maintained: last updated July 14, 2026, adding GPT-5.6 Sol, Muse Spark 1.1, Grok 4.5
  • Current frontier coverage: Claude Fable 5 (top), GPT-5.5/5.6, Gemini 3.5 Flash, DeepSeek V4 Pro, Kimi K2.6, Grok 4.3-4.5, Qwen 3.6-3.7

Data-reality check — NOT INGESTIBLE (yet)

Checked the repo on 2026-07-20: the root contains only images/, prompts_wc/, stories_wc/ directories and README.md. Leaderboard scores are rendered as tables inside the README — no machine-readable export exists. The published bundle contains prompt and story text files, not scores.

What the owner would need to publish

A CSV or JSON file at the repo root (e.g. results.csv) with at least: model, score (the comparison score), and ideally rank. Something like:

model,score,rank
claude-fable-5-high,3.3,1
gpt-5-6-sol-high,3.2,2
...

Once that file is published and accessible via raw GitHub URL, this bench is a strong build candidate.

Metadata

Metadata

Assignees

No one assigned

    Labels

    candidateneeds-dataInteresting bench, but no machine-readable / current data to ingest yet

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions