A curated list of papers, tools, datasets, benchmarks, and standards for building, evaluating, and auditing reliable AI agents.
-
Updated
Jun 30, 2026 - HTML
A curated list of papers, tools, datasets, benchmarks, and standards for building, evaluating, and auditing reliable AI agents.
SAGE: self-correcting autonomous research via Multi-Hypothesis Failure Attribution (MHFA). Diagnose many causes of a failed experiment, verify the most critical, and route the fix to the right level.
Observability and a transparent, auditable agent graph for Massive Intelligence (IM) swarms: OpenTelemetry-compatible span tracing, incremental drift detection, failure attribution, loop detection, thirteen alert channels, and SHA-256-chained, Ed25519-signed trace logs. Zero dependencies, node-free core, TypeScript-first.
Telemetry-grounded, calibrated failure attribution for agent oversight (OTel GenAI + Who&When).
Failure attribution for agent pipelines — find which span caused the failure and what kind of fix it needs.
When the agent breaks, which layer dropped the ball? Operator-facing failure-attribution taxonomy for AI agent estates: ten-way dictionary, MAST + AgentRx crosswalks, postmortem template. Part of the Spine catalog.
Record, replay, and automatically attribute failures in multi-agent LLM systems — deterministic replay, counterfactual + learned root-cause attribution, and verified fixes. Offline by default.
Add a description, image, and links to the failure-attribution topic page so that developers can more easily learn about it.
To associate your repository with the failure-attribution topic, visit your repo's landing page and select "manage topics."