Add ClawBench to agent benchmark resources - #64
Conversation
|
Thanks — and this is the version with the correct link. I confirmed Two changes before merge:
Also for the record: you're a maintainer of ClawBench, which you disclosed on #54 but not here — please add that disclosure. It doesn't change the verdict, but the list depends on affiliations being visible. Once the section moves, I'll merge this and close #54 as superseded. |
|
Small update: I've already closed #54 in favour of this PR, so there's nothing pending on your side there — my earlier comment made it sound conditional. The two things before merge are unchanged: move the entry from §6 to §9, and add the maintainership disclosure. |
|
Independent verification check before this merges. Paper numbers confirmed against arXiv 2604.08523: 153 tasks, 144 live platforms, 15 categories — all verbatim from the abstract. TIGER-AI-Lab/ClawBench is at 538★ with a push on 2026-07-28, so traction and activity are real. One addition worth considering for the annotation: the abstract gives a concrete anchor — "Claude Sonnet 4.6 achieves only 33.3%" — which tells a reader immediately that this benchmark isn't saturated and calibrates how hard the tasks actually are. That kind of number is exactly what a practitioner wants before deciding whether to run it. The two open items from the earlier review still need to land:
Recommendation: needs changes. Once the section moves, the disambiguation tag is in, and the disclosure is added — this is a straightforward merge. The benchmark clears the bar: real production websites, a genuine interception mechanism, active development, and a published arXiv paper with reproducible task counts. |
What changed
Add ClawBench to the benchmark-integrity section.
Why
ClawBench is a directly relevant real-world agent benchmark: it evaluates 153 everyday browser tasks across 144 live production websites and uses submission interception to prevent side effects while retaining realistic workflows and detailed traces. This complements the section's focus on trustworthy, reproducible agent evaluations.
Validation
git diff --check