Pinned Loading
-
dafny-eval
dafny-eval PublicSoundness-hardened benchmark for LLM Dafny invariant synthesis — semantic-integrity gating, capability ladder, prompt ablation, and an agentic verifier-feedback repair loop.
Python
-
evalaware
evalaware PublicAn audit of linear eval-awareness probes on real transcripts: held-out-source generalisation and reliability at a realistic base rate (Qwen3-4B/8B, jjpn2/eval_awareness)
Python
-
scorer-integrity
scorer-integrity PublicCan the scheming detector be trusted? Calibrating LLM judges against mechanical ground truth in Inspect evals.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.
