Skip to content

Add ClawBench to agent benchmark resources - #64

Open
reacher-z wants to merge 1 commit into
benchflow-ai:mainfrom
reacher-z:agent/add-clawbench
Open

Add ClawBench to agent benchmark resources#64
reacher-z wants to merge 1 commit into
benchflow-ai:mainfrom
reacher-z:agent/add-clawbench

Conversation

@reacher-z

Copy link
Copy Markdown

What changed

Add ClawBench to the benchmark-integrity section.

Why

ClawBench is a directly relevant real-world agent benchmark: it evaluates 153 everyday browser tasks across 144 live production websites and uses submission interception to prevent side effects while retaining realistic workflows and detailed traces. This complements the section's focus on trustworthy, reproducible agent evaluations.

Validation

  • git diff --check
  • Verified the arXiv and repository links.

@xdotli

xdotli commented Jul 30, 2026

Copy link
Copy Markdown
Member

Thanks — and this is the version with the correct link. I confirmed github.com/reacher-z/ClawBench 301-redirects to TIGER-AI-Lab/ClawBench, so this PR has the canonical URL while #54 does not. The paper checks out too (arXiv 2604.08523 — 153 tasks / 144 platforms / 15 categories, verbatim in the abstract), and 538★ with active pushes is real traction.

Two changes before merge:

  1. Section. This inserts into §6 (benchmark vs. eval, and benchmark integrity — contamination, saturation, leaderboard gaming). ClawBench is a browser-agent benchmark, so it belongs in §9 · Agent-specific evaluation, which is where Add ClawBench to agent-specific evaluation #54 placed it. Could you move it there?

  2. Disambiguation. §9 already contains ClawsBench (README:352), which is BenchFlow's own, unrelated benchmark. The two names differ by one letter and would sit a few lines apart. Please add a short ⚠️ distinct from ClawsBench (BenchFlow) to the entry.

Also for the record: you're a maintainer of ClawBench, which you disclosed on #54 but not here — please add that disclosure. It doesn't change the verdict, but the list depends on affiliations being visible.

Once the section moves, I'll merge this and close #54 as superseded.

@xdotli

xdotli commented Jul 30, 2026

Copy link
Copy Markdown
Member

Small update: I've already closed #54 in favour of this PR, so there's nothing pending on your side there — my earlier comment made it sound conditional.

The two things before merge are unchanged: move the entry from §6 to §9, and add the maintainership disclosure.

@github-actions

Copy link
Copy Markdown
Contributor

Independent verification check before this merges.

Paper numbers confirmed against arXiv 2604.08523: 153 tasks, 144 live platforms, 15 categories — all verbatim from the abstract. TIGER-AI-Lab/ClawBench is at 538★ with a push on 2026-07-28, so traction and activity are real.

One addition worth considering for the annotation: the abstract gives a concrete anchor — "Claude Sonnet 4.6 achieves only 33.3%" — which tells a reader immediately that this benchmark isn't saturated and calibrates how hard the tasks actually are. That kind of number is exactly what a practitioner wants before deciding whether to run it.

The two open items from the earlier review still need to land:

  1. Section: The diff currently inserts into §6 (benchmark integrity). ClawBench is a browser-agent benchmark and belongs in §9 · Agent-specific evaluation.
  2. Disambiguation: §9 already carries ClawsBench (BenchFlow's browser-automation benchmark, README:352) — one letter apart. Please add ⚠️ distinct from ClawsBench (BenchFlow) to the entry so readers don't confuse the two.
  3. Disclosure: Please add a maintainer note to the PR body (as you did on Add ClawBench to agent-specific evaluation #54).

Recommendation: needs changes. Once the section moves, the disambiguation tag is in, and the disclosure is added — this is a straightforward merge. The benchmark clears the bar: real production websites, a genuine interception mechanism, active development, and a published arXiv paper with reproducible task counts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants