Skip to content

Repository files navigation

ModelGauntlet

Judge-ready CI

Live app: https://modelgauntlet.vercel.app

Break your AI before users do.

ModelGauntlet stress-tests a structured AI feature across open models and returns a reproducible SHIP, FIX, or BLOCK decision before a developer deploys it.

Built solo for Impact Forge 2026 — General Innovation.

ModelGauntlet release gate

The problem

A prompt can look excellent on a happy path and still fail on missing evidence, output format, policy boundaries, or prompt injection. Solo developers rarely have time to assemble a dataset, integrate several model vendors, and build an evaluation harness before a demo or release.

ModelGauntlet turns that uncertain moment into one short workflow:

Task contract → identical cross-model run → exact deterministic checks → SHIP / FIX / BLOCK

What makes this different

  • A release decision, not another chat response. The primary output is an actionable gate with an explicit threshold.
  • AI does not grade AI. Featherless generates bounded adversarial cases and candidate responses; deterministic TypeScript assigns every core score.
  • Human holdouts remain untouched. Generated tests expand coverage but cannot replace the critical release boundary.
  • Every failure is reproducible. Judges can open the exact input, raw output, normalization step, assertion, latency, and token usage.
  • The sponsor is structural. Featherless exposes multiple open models through one OpenAI-compatible interface, so the same suite can be compared without separate provider integrations or local GPUs.
  • The demo survives dependencies. A persistent, clearly labelled seeded snapshot demonstrates the product workflow without pretending to be live inference.

Judge demo — 60 seconds

  1. Open modelgauntlet.vercel.app and keep Seeded proof selected.
  2. Click Run baseline gauntlet.
  3. ModelGauntlet returns BLOCK because every model fails at least one critical human holdout.
  4. Open Expired return window → Exact trace to see the raw APPROVE output fail the expected DENY assertion.
  5. Click Apply hardening & rerun.
  6. The best candidate reaches SHIP; the report shows the baseline delta, untouched-holdout result, latency, tokens, and remaining Qwen failure.

The seeded numbers are fixture data, visibly labelled in the interface. Switch to Live Featherless on a configured deployment to execute actual API requests.

Architecture

flowchart LR
    U["Developer / judge"] --> W["Next.js experiment workspace"]
    W --> G["POST /api/generate-cases"]
    G --> F1["Featherless generator model"]
    F1 --> V["Zod bounded-case validation"]
    V --> W

    W -->|"max 3 concurrent"| E["POST /api/evaluate<br/>one model × one case"]
    E --> F2["Featherless open models"]
    F2 --> N["Conservative JSON normalization"]
    N --> D["Ajv JSON Schema + exact TypeScript assertions"]
    D --> R["Evidence report<br/>SHIP / FIX / BLOCK"]
    R --> U

    S["Labelled seeded snapshot"] -. "dependency fallback" .-> W
Loading

Trust boundary

AI is allowed to Deterministic code owns
Propose up to four adversarial inputs Request and generated-case validation
Execute the structured task under test JSON parsing and normalization disclosure
Produce candidate JSON responses JSON Schema, exact path, phrase, and latency checks
Fail partially without erasing other models Concurrency, timeouts, aggregation, and final verdict

Verdict policy:

  • SHIP: every critical human-holdout assertion passes and total pass rate is at least 90%.
  • FIX: every critical human-holdout assertion passes and total pass rate is at least 70%.
  • BLOCK: a critical human-holdout assertion fails or total pass rate is below 70%.

Reliability evidence

Current local verification:

  • 23 automated tests across JSON recovery, assertion types, schema-dialect behavior, verdict boundaries, generated-risk normalization, seeded-fixture recomputation, Featherless response/error mapping, and no-key route behavior.
  • ESLint clean.
  • TypeScript strict check clean.
  • Next.js production build clean.
  • Browser verified at 1280 px and 390 px widths with no framework overlay, console error, or horizontal overflow.
  • Exact failure drawer and baseline-to-hardened seeded path verified in a real browser.

Public production proof collected on 2026-08-15:

  • Qwen 2.5 7B generated two adversarial cases that passed the bounded Zod contract (497 total tokens).
  • A hardened eligible-refund holdout passed all five deterministic assertions in 2,865 ms (233 total tokens).
  • A hardened injection holdout returned valid JSON but passed only four of five assertions in 1,926 ms (213 total tokens), demonstrating that the live gate catches a plausible model response rather than manufacturing a success.

These are individual verification runs, not latency benchmarks. Seeded fixture values remain visibly labelled and are never presented as live measurements.

Run locally

Requirements: Node.js 20.9+ and a Featherless API key for live inference.

git clone https://github.com/sgoel2be24-cyber/modelgauntlet.git
cd modelgauntlet
npm install
cp .env.example .env.local
npm run dev

Add your server-only key to .env.local:

FEATHERLESS_API_KEY=your_key_here
FEATHERLESS_GENERATOR_MODEL=Qwen/Qwen2.5-7B-Instruct
NEXT_PUBLIC_APP_URL=http://localhost:3000

Open http://localhost:3000. Snapshot mode works without a key. Live evaluation and adversarial regeneration enable automatically when the server detects the key.

The key is read only in server route handlers and is never returned by /api/health or prefixed with NEXT_PUBLIC_.

Verify

npm test
npm run lint
npx tsc --noEmit
npm run build

Project structure

app/api/                 Featherless evaluation, generation, and safe health routes
components/              Judge-facing workbench, report, comparison, and trace UI
lib/domain/              Typed contract and Zod request schemas
lib/evaluation/          JSON normalization, assertions, prompts, and verdict rules
lib/featherless/         Server-only OpenAI-compatible Featherless client
lib/fixtures/            Seeded support contract and labelled fallback report
lib/orchestration/       Bounded client concurrency
tests/                   Deterministic engine, route, fixture, and client tests
docs/hackathon-build/    Scope, PRD, technical spec, checklist, and build notes

Failure handling

  • Invalid request contracts return structured validation errors.
  • Missing credentials disable live controls while preserving snapshot mode.
  • Rate limits, timeouts, and upstream failures become inspectable per-case failures.
  • A failed model request does not erase completed results from other models.
  • Markdown-wrapped or prose-wrapped JSON is recovered only through conservative strategies, and the normalization is recorded in the trace.
  • Editing a prompt or test suite invalidates old UI evidence instead of silently reusing it.

Explicit limitations

  • MVP supports structured text-to-JSON tasks, not arbitrary code or external API execution.
  • Generated adversarial cases expand coverage; they do not certify universal safety.
  • No database, authentication, teams, production proxy, fine-tuning, or subjective LLM judge.
  • Model IDs and availability can change; verify the selected IDs in the Featherless catalog before the final live demo.

Documentation and citations

Hackathon materials

License

MIT — see LICENSE.

About

Break your AI before users do — deterministic release evidence across open models.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages