Multi-tier framework for evaluating AI agent skills with quality gates, semantic overlap detection, synthetic evaluation dataset generation, and live agent evaluation that measures how skills affect agent behavior.
benchmark skills evaluation security-scanner codex evaluate skill-evaluation agentic-ai agent-skills agent-security claude-code agent-evaluation skill-eval skill-evaluator skill-evals
-
Updated
Sep 1, 2026 - Python