Live app: https://modelgauntlet.vercel.app
Break your AI before users do.
ModelGauntlet stress-tests a structured AI feature across open models and returns a reproducible SHIP, FIX, or BLOCK decision before a developer deploys it.
Built solo for Impact Forge 2026 — General Innovation.
A prompt can look excellent on a happy path and still fail on missing evidence, output format, policy boundaries, or prompt injection. Solo developers rarely have time to assemble a dataset, integrate several model vendors, and build an evaluation harness before a demo or release.
ModelGauntlet turns that uncertain moment into one short workflow:
Task contract → identical cross-model run → exact deterministic checks → SHIP / FIX / BLOCK
- A release decision, not another chat response. The primary output is an actionable gate with an explicit threshold.
- AI does not grade AI. Featherless generates bounded adversarial cases and candidate responses; deterministic TypeScript assigns every core score.
- Human holdouts remain untouched. Generated tests expand coverage but cannot replace the critical release boundary.
- Every failure is reproducible. Judges can open the exact input, raw output, normalization step, assertion, latency, and token usage.
- The sponsor is structural. Featherless exposes multiple open models through one OpenAI-compatible interface, so the same suite can be compared without separate provider integrations or local GPUs.
- The demo survives dependencies. A persistent, clearly labelled seeded snapshot demonstrates the product workflow without pretending to be live inference.
- Open modelgauntlet.vercel.app and keep Seeded proof selected.
- Click Run baseline gauntlet.
- ModelGauntlet returns
BLOCKbecause every model fails at least one critical human holdout. - Open Expired return window → Exact trace to see the raw
APPROVEoutput fail the expectedDENYassertion. - Click Apply hardening & rerun.
- The best candidate reaches
SHIP; the report shows the baseline delta, untouched-holdout result, latency, tokens, and remaining Qwen failure.
The seeded numbers are fixture data, visibly labelled in the interface. Switch to Live Featherless on a configured deployment to execute actual API requests.
flowchart LR
U["Developer / judge"] --> W["Next.js experiment workspace"]
W --> G["POST /api/generate-cases"]
G --> F1["Featherless generator model"]
F1 --> V["Zod bounded-case validation"]
V --> W
W -->|"max 3 concurrent"| E["POST /api/evaluate<br/>one model × one case"]
E --> F2["Featherless open models"]
F2 --> N["Conservative JSON normalization"]
N --> D["Ajv JSON Schema + exact TypeScript assertions"]
D --> R["Evidence report<br/>SHIP / FIX / BLOCK"]
R --> U
S["Labelled seeded snapshot"] -. "dependency fallback" .-> W
| AI is allowed to | Deterministic code owns |
|---|---|
| Propose up to four adversarial inputs | Request and generated-case validation |
| Execute the structured task under test | JSON parsing and normalization disclosure |
| Produce candidate JSON responses | JSON Schema, exact path, phrase, and latency checks |
| Fail partially without erasing other models | Concurrency, timeouts, aggregation, and final verdict |
Verdict policy:
SHIP: every critical human-holdout assertion passes and total pass rate is at least 90%.FIX: every critical human-holdout assertion passes and total pass rate is at least 70%.BLOCK: a critical human-holdout assertion fails or total pass rate is below 70%.
Current local verification:
- 23 automated tests across JSON recovery, assertion types, schema-dialect behavior, verdict boundaries, generated-risk normalization, seeded-fixture recomputation, Featherless response/error mapping, and no-key route behavior.
- ESLint clean.
- TypeScript strict check clean.
- Next.js production build clean.
- Browser verified at 1280 px and 390 px widths with no framework overlay, console error, or horizontal overflow.
- Exact failure drawer and baseline-to-hardened seeded path verified in a real browser.
Public production proof collected on 2026-08-15:
- Qwen 2.5 7B generated two adversarial cases that passed the bounded Zod contract (497 total tokens).
- A hardened eligible-refund holdout passed all five deterministic assertions in 2,865 ms (233 total tokens).
- A hardened injection holdout returned valid JSON but passed only four of five assertions in 1,926 ms (213 total tokens), demonstrating that the live gate catches a plausible model response rather than manufacturing a success.
These are individual verification runs, not latency benchmarks. Seeded fixture values remain visibly labelled and are never presented as live measurements.
Requirements: Node.js 20.9+ and a Featherless API key for live inference.
git clone https://github.com/sgoel2be24-cyber/modelgauntlet.git
cd modelgauntlet
npm install
cp .env.example .env.local
npm run devAdd your server-only key to .env.local:
FEATHERLESS_API_KEY=your_key_here
FEATHERLESS_GENERATOR_MODEL=Qwen/Qwen2.5-7B-Instruct
NEXT_PUBLIC_APP_URL=http://localhost:3000
Open http://localhost:3000. Snapshot mode works without a key. Live evaluation and adversarial regeneration enable automatically when the server detects the key.
The key is read only in server route handlers and is never returned by /api/health or prefixed with NEXT_PUBLIC_.
npm test
npm run lint
npx tsc --noEmit
npm run buildapp/api/ Featherless evaluation, generation, and safe health routes
components/ Judge-facing workbench, report, comparison, and trace UI
lib/domain/ Typed contract and Zod request schemas
lib/evaluation/ JSON normalization, assertions, prompts, and verdict rules
lib/featherless/ Server-only OpenAI-compatible Featherless client
lib/fixtures/ Seeded support contract and labelled fallback report
lib/orchestration/ Bounded client concurrency
tests/ Deterministic engine, route, fixture, and client tests
docs/hackathon-build/ Scope, PRD, technical spec, checklist, and build notes
- Invalid request contracts return structured validation errors.
- Missing credentials disable live controls while preserving snapshot mode.
- Rate limits, timeouts, and upstream failures become inspectable per-case failures.
- A failed model request does not erase completed results from other models.
- Markdown-wrapped or prose-wrapped JSON is recovered only through conservative strategies, and the normalization is recorded in the trace.
- Editing a prompt or test suite invalidates old UI evidence instead of silently reusing it.
- MVP supports structured text-to-JSON tasks, not arbitrary code or external API execution.
- Generated adversarial cases expand coverage; they do not certify universal safety.
- No database, authentication, teams, production proxy, fine-tuning, or subjective LLM judge.
- Model IDs and availability can change; verify the selected IDs in the Featherless catalog before the final live demo.
- Featherless quickstart
- Featherless chat-completions endpoint
- Featherless API overview and attribution headers
- Next.js App Router
- Next.js Route Handlers
- JSON Schema 2020-12
- Ajv JSON Schema validator
- Zod validation
- Product requirements
- Technical specification
- Build checklist
- Verification record
- Three-minute demo script
- Devpost technical overview
MIT — see LICENSE.
