Track 17,170+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
-
Updated
Sep 27, 2026 - Python
Track 17,170+ AI benchmark, eval, dataset, and data-quality records from 37 public sources, with linked evidence and daily updates.
MyPhoneBench: Do Phone-Use Agents Respect Your Privacy?
20 runnable LLM agent design patterns in Python, benchmarked, traced, offline-compatible. ReAct, multi-agent, constitutional AI, and more.
The same AI agent pipeline built in Mastra and LangChain. Runs in parallel, measures everything.
splyntra
A Claude Agent SDK security benchmark project
SnorkelAI Terminus2 task suite with deterministic evaluators, Dockerized environments, and reproducible agent-benchmarking primitives.
Deterministic multi-lane evaluation domain for SnorkelAI Terminus3. Cockpit, lane orchestration, agent benchmarking, reasoning-drift detection.
Correctness-gated measurement of mutation, edit payload, and semantic revisit across agent trajectories.
To associate your repository with the agent-benchmarking topic, visit your repo's landing page and select "manage topics."