Public evidence for AIC coding-agent evaluations: frozen patches, canonical verifier scores, integrity hashes, cost, timing, and task context.
benchmark evaluation software-engineering reproducibility ai-agents qwen llm-evaluation coding-agents deepswe deep-swe
-
Updated
Sep 16, 2026 - PowerShell