Reproducible coding-model benchmark with auditable maintenance tasks, candidate patches, execution records, evaluator reports, hidden-test results, and comparative code-quality reviews.
javascript benchmarking reproducible-research opencode software-engineering code-generation code-quality codex model-evaluation model-comparison test-fixtures llm llm-evaluation coding-agents llm-benchmark ai-model-benchmark
-
Updated
Aug 31, 2026 - JavaScript