Auditable DeepSWE v1.1 evaluation of Agent Claim Network: coding-harness pass rate and cross-agent claim reuse|ACN 在 DeepSWE v1.1 上的可复现评测:coding harness 通过率与跨 agent 的 claim 复用收益。
-
Updated
Sep 17, 2026 - Python
Auditable DeepSWE v1.1 evaluation of Agent Claim Network: coding-harness pass rate and cross-agent claim reuse|ACN 在 DeepSWE v1.1 上的可复现评测:coding harness 通过率与跨 agent 的 claim 复用收益。
Public evidence for AIC coding-agent evaluations: frozen patches, canonical verifier scores, integrity hashes, cost, timing, and task context.
DeepSWE v1.1: Perfect score (113/113) — all tasks solved with reward=1.0
Reproducible analysis of reasoning-effort saturation on DeepSWE v1.1
Forensic audit scanner for SWE-bench & DeepSWE v1.1 containers detecting .git reflog leakage and test harness spoofing.
DeepSWE Launcher ⚡ lightweight DeepSWE benchmark launcher: batching, resume, multi-model profiles, live dashboard (terminal + WebUI), Excel reports
Daily mirror of the DeepSWE live leaderboard with TrueIQ benchmarking analysis.
To associate your repository with the deepswe topic, visit your repo's landing page and select "manage topics."