Skip to content

Add ClawBench to agent-specific evaluation - #54

Closed
reacher-z wants to merge 2 commits into
benchflow-ai:mainfrom
reacher-z:add-clawbench-agent-eval
Closed

Add ClawBench to agent-specific evaluation#54
reacher-z wants to merge 2 commits into
benchflow-ai:mainfrom
reacher-z:add-clawbench-agent-eval

Conversation

@reacher-z

@reacher-z reacher-z commented Jul 26, 2026

Copy link
Copy Markdown

Summary

Adds one annotated ClawBench entry to section 9, beside live-web and trajectory-evaluation benchmarks.

Why it clears the inclusion bar

  • open-source code and an accompanying arXiv paper
  • 153 consequential live-site tasks spanning 144 websites and 15 life domains
  • isolated task containers and support for multiple agent harnesses
  • replay, action, network, and message traces that support both outcome and trajectory analysis

The entry follows the repository’s Name — Org — URL · type — one-line why format and marks this 2026 work as new.

Disclosure: I am a ClawBench maintainer submitting this entry for consideration.

Project links: Website · GitHub · Paper

@xdotli

xdotli commented Jul 30, 2026

Copy link
Copy Markdown
Member

Closing this one in favor of #64, which covers the same paper — but to be clear about which parts of each were right, since it's a split decision:

#54 got the section right. §9 · Agent-specific evaluation is where ClawBench belongs; #64 placed it in §6 and I've asked for it to move here.

#64 got the URL right. github.com/reacher-z/ClawBench 301-redirects to TIGER-AI-Lab/ClawBench, so the link in this PR is a stale personal alias.

One correction worth carrying over: this entry says the tasks run in "isolated task containers," but the abstract makes the opposite point — that ClawBench operates on production websites rather than offline sandboxes, with a submission-interception layer as the safeguard. That distinction is the benchmark's main selling point, so it's worth stating accurately in #64.

Also: this is the fifth PR for ClawBench (#56, #59, #60 self-closed, plus #54 and #64 both open simultaneously). All good — but going forward, one open PR per resource makes review much easier, and #64 is the one to iterate on.

@xdotli xdotli closed this Jul 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants