ORDEAL
Test what an agent does, observe what it did, and turn failures into lasting regression tests.
Ordeal tests what an AI agent actually does, not what it says it did. A Python simulator runs your agent against stateful, tool-calling scenarios and checks the resulting trajectory and final state against deterministic assertions — catching things a plain "did the output look right?" check misses, like an agent claiming to complete an action it never took. A self-hosted server and web console turn captured production traces into versioned regression suites.
Everything runs on your own infrastructure — no hosted account, no Kafka, no proprietary service. A TypeScript client is included for non-Python agents.
This is an evaluation/pilot release, not a certified production system. See Known limitations and the capability matrix before relying on any specific feature.
The web console after comparing four models on the same suite via ordeal-behavior experiment:
Requires Python 3.10+. Not yet published to PyPI or npm — install from this source tree.
python -m venv .venv && source .venv/bin/activate
python -m pip install -e '.[server,otel]'
ordeal-server init --name 'My organization'
ordeal-server serveOpen http://127.0.0.1:8080 and sign in with the token init prints once. For local behavior testing only, server dependencies aren't required:
ordeal-behavior init
ordeal-behavior run examples/enterprise/suite.py --json report.json --junit report.xmlPoint LLMAgent at a model and a World, and it drives the full tool-calling loop itself — deriving the function-calling schema from the world's own tools, rate-limiting and retrying requests, recovering from a malformed tool call, and tracking token/cost usage.
from ordeal_agent import LLMAgent, Scenario, StateEquals, Suite, ToolCalled, ToolOrder, World, simulated
def lookup_order(args, ctx):
return ctx.world.get("orders", {}).get(args["order_id"], {"status": "not_found"})
def refund_order(args, ctx):
orders = ctx.world.get("orders", {})
order = orders.get(args["order_id"])
if order is None or order["status"] != "paid":
return {"ok": False}
ctx.world.set("orders", {**orders, args["order_id"]: {**order, "status": "refunded"}})
return {"ok": True}
world = World(
"support",
initial_state={"orders": {"A-100": {"status": "paid", "amount": 42}}},
tools=[
simulated("lookup_order", lookup_order, description="Look up an order",
input_schema={"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}),
simulated("refund_order", refund_order, description="Refund an eligible order",
input_schema={"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}),
],
)
agent = LLMAgent(
world=world,
model="google/gemini-2.5-flash-lite", # any OpenAI-compatible model id
base_url="https://openrouter.ai/api/v1", # OpenAI, OpenRouter, vLLM, Ollama, ...
api_key_env="OPENROUTER_API_KEY",
)
suite = Suite.of("refund", Scenario(
"refund-paid-order",
"A customer wants a refund on order A-100. Look it up and refund it if eligible.",
world,
assertions=[ToolCalled("lookup_order"), ToolCalled("refund_order"),
ToolOrder("lookup_order", "refund_order"),
StateEquals("orders", {"A-100": {"status": "refunded", "amount": 42}})],
))ToolCalled and StateEquals together catch a model that says "I've refunded your order" without ever calling refund_order — a failure plain output-matching misses entirely. That's the case that motivated LLMAgent; see ordeal_tests/test_openrouter_agent.py and ordeal_tests/test_enterprise_experiment.py for the full suites behind the screenshots above, and pass a list of Variant(model, LLMAgent(...)) to ordeal-behavior experiment to compare models by pass rate, cost, and latency.
| Capture production traces | PlatformClient(...).trace(...) spans, or OTLP/HTTP+gRPC — see docs/SDK_GUIDE.md |
| Turn a trace into a regression test | Console → dataset → client.replay_dataset(...) with your own assertions |
| Run suites on a private worker | ordeal-server runner --allow-suite ... — a real (non-sandboxed) child process; see docs/enterprise/DEPLOYMENT.md |
| Framework guide, API reference, security, operations | docs/enterprise/ |
- All eleven framework adapter shims (LangChain, CrewAI, AutoGen, LlamaIndex, Strands, Google ADK, OpenAI Agents SDK, pydantic-ai, smolagents, MCP, plus
LLMAgent) now pass CI against the real installed framework packages — three (LangChain, AutoGen, smolagents) had real bugs that silently dropped tool-call arguments or failed schema validation, fixed and verified. OnlyLLMAgenthas actually driven a live model end-to-end, though; the others verify the tool-wrapping shim, not a real LLM deciding to call it through that framework. - The Docker runner backend has been run end-to-end against a live container daemon and escape-tested against network egress, filesystem writes, capability/setuid abuse, the classic
cgroup release_agentescape, tmpfs exec, and fork bombs — all blocked. One real bug found along the way: on an SELinux-enforcing host (Fedora/RHEL-family), the bind-mounted workspace was denied at the MAC layer regardless of Unix permissions; fixed with a per-mount relabel (-v ...:ro,z) rather than disabling confinement for the whole container. No kernel-exploit-based escape attempt or independent security review has been done.trusted-processmode remains explicitly not a sandbox. - Browser/console tests run with native Chromium navigation in CI (no transport override) and were independently re-verified locally the same way.
- No independent penetration test, SOC 2/ISO process, or enterprise-scale (Postgres/S3/HA) validation — see
docs/enterprise/COMMERCIAL_READINESS.mdbefore selling this as a hosted service.
Full itemized status: docs/enterprise/CAPABILITY_MATRIX.md · Test report: docs/TESTING.md

