Add ClawBench to agent-specific evaluation - #54
Conversation
|
Closing this one in favor of #64, which covers the same paper — but to be clear about which parts of each were right, since it's a split decision: #54 got the section right. §9 · Agent-specific evaluation is where ClawBench belongs; #64 placed it in §6 and I've asked for it to move here. #64 got the URL right. One correction worth carrying over: this entry says the tasks run in "isolated task containers," but the abstract makes the opposite point — that ClawBench operates on production websites rather than offline sandboxes, with a submission-interception layer as the safeguard. That distinction is the benchmark's main selling point, so it's worth stating accurately in #64. Also: this is the fifth PR for ClawBench (#56, #59, #60 self-closed, plus #54 and #64 both open simultaneously). All good — but going forward, one open PR per resource makes review much easier, and #64 is the one to iterate on. |
Summary
Adds one annotated ClawBench entry to section 9, beside live-web and trajectory-evaluation benchmarks.
Why it clears the inclusion bar
The entry follows the repository’s Name — Org — URL · type — one-line why format and marks this 2026 work as new.
Disclosure: I am a ClawBench maintainer submitting this entry for consideration.
Project links: Website · GitHub · Paper