Threshold testing is blind to agent regression. agent-eval runs your agent 50x on version A and B and gives you a p-value on whether behavior actually shifted, not just whether one run looked different.
python testing statistics ci-cd p-value regression-testing effect-size langchain llm-evaluation llm-agents crewai langgraph llm-observability llm-testing agent-evaluation openai-agents-sdk agent-testing mann-whitney-u bootstrap-confidence-interval promptfoo-alternative
-
Updated
Jul 22, 2026 - Python