Skip to content
tendaigomoPublic

About

The Taravangian Test for AI: catch silent model degradation before your users do. SPC-based reasoning-quality monitoring for Claude, GPT, Gemini, and Grok.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

NerfWatch, the cardiac monitor for AI models

Go Report Card License: MIT Go Version

The Taravangian Test for AI.
Catch silent model degradation before your users do.

In Brandon Sanderson's Stormlight Archive, King Taravangian wakes each day not knowing whether he'll be a genius or barely able to function. So every morning he takes the same test, measures the result, and only then decides what he can be trusted to do that day.

AI models have the same condition. Providers retrain, requantize, and reroute without telling you, so the model you built against is not guaranteed to be the model serving your users right now. The only way to know which one woke up is to test it. Same test, every morning, measured.

NerfWatch is that test, run as a monitor. It probes your models on a schedule, pulls behavioral vital signs out of every response, and feeds the numbers through the same statistical process control charts hospitals use to watch patients. When a chart fires, you get an alert that names the signal, the severity, and the moment things changed.

What NerfWatch is

A standalone Go binary that watches AI models from the outside. No SDK in your code, no proxy in your request path. You give it API keys; it sends parameterized probes on its own schedule, grades every answer with a different provider's model, and charts five behavioral signals over time: thinking tokens, output inflation, latency, early stopping, connector density.

This is not a one-vendor problem. OpenAI users spent December 2023 documenting GPT-4's "laziness" until OpenAI acknowledged it and shipped a fix. In April 2025 an update made GPT-4o sycophantic enough that OpenAI rolled it back and published a postmortem. And in early 2026, Claude Opus silently lost most of its reasoning depth: thinking tokens fell 73%, reads-per-edit fell from 6.6 to 2.0, and the proof came not from the provider but from one engineer who hand-analyzed 6,852 sessions. Three incidents, two vendors, one pattern. The next one can come from any of them, which is why NerfWatch monitors Anthropic, OpenAI, Google, and xAI models with identical instruments and no favorites. It exists so the next proof takes one to three probe cycles, not six thousand sessions.

The frontier has moved on since those incidents, and visibility has gotten worse, not better. Fable 5 never returns its raw chain of thought; you get a summary or nothing. GPT-5.6 Sol reports reasoning as a bare token count. Gemini 3.1 Pro and Grok 4.5 still expose reasoning text today, with no guarantee about tomorrow. The stronger the model, the less of its thinking you are allowed to see. What's left is the outside view: measure behavior the way a doctor reads vitals off a patient who can't tell you how they feel. That is the whole product, and it works the same on all four providers.

Principles

These follow from the Taravangian condition. You can't know which model woke up, so the design assumes you have to find out every cycle:

  • Test every cycle, not once. A model that passed your evaluation at launch can quietly become a different model tomorrow. NerfWatch re-takes the test on a schedule and never assumes yesterday's result still holds.
  • A degraded model can't grade itself. The bent ruler cannot measure its own bend, so grading is always cross-provider. No model scores its own output.
  • Measure each dimension, not an average. The Opus collapse was about 4% as a single aggregate score, dismissable noise, yet a 25-point drop in chain-of-thought depth. NerfWatch tracks logic, math, factual, and code separately so a collapse in one cannot hide behind health in the others.
  • Statistics, not thresholds. EWMA, CUSUM, and BOCPD catch the degradation you didn't anticipate, not just the failures you remembered to write a rule for.
  • Private, parameterized probes. Any published test set is eventually trained on. NerfWatch's probes vary per run from a stored seed: reproducible, never memorizable.
  • Derived metrics only, local-first. NerfWatch records statistical signals, never raw model output, in a SQLite database on your machine. No cloud, no telemetry.

What NerfWatch catches

Scenario Layer Detection Time to alert
Opus-style reasoning collapse (73% thinking-token drop) Layer 1 MEWS CRITICAL 1-3 runs
Gradual CoT quality erosion (10% over 2 weeks) Layer 2 + CUSUM DRIFT Days, not weeks
Provider hides thinking tokens Layer 1 signal loss WARNING Immediate
Code generation correctness regression Layer 2 code probes WARNING/CRITICAL 1-3 runs
Model upgrade changes behavior Layer 1 + BOCPD Changepoint Exact moment identified

What the numbers mean

Every NerfWatch alert names the instrument that fired. None of these are invented for this project; they come from clinical wards and factory floors, where the cost of a missed shift is measured in lives and recalls. Here is each one, what it measures, and why it earns its place.

Thinking tokens. How much reasoning the model did before answering. Counted from reasoning text where the provider exposes it, from usage metadata where it doesn't. This was the first visible symptom of the Opus collapse: a model that used to spend 1,200 tokens working a logic probe started spending 300. Whatever the final answer says, that is a different model. Alerts fire on drops; a model thinking more is not an emergency.

Inflation ratio. Output tokens relative to input. Degrading models often get wordier while getting shallower; verbosity fills the space reasoning used to occupy. During the Opus incident, output volume rose while thinking fell. Watching both catches the trade. Alerts fire on rises.

Latency profile. Milliseconds per output token. A sudden change here usually means the serving stack changed underneath you: different hardware, different quantization, different routing. That is not degradation by itself, which is why latency votes in the composite score instead of raising alarms alone.

Early stopping. How often the model interrupts its own work ("Shall I continue?", answers that trail off). A model that increasingly asks permission instead of finishing is failing in a way accuracy tests miss, because the answers it does complete may still be correct.

Connector density. Logical connectors (therefore, because, however) per 100 words of reasoning text. Sound reasoning is argument, and argument leaves scaffolding. When chain-of-thought decays from argument into narration, the connectors thin out before the answers go wrong, so this is often the earliest signal to move. Not measurable on OpenAI models, which don't expose reasoning text; the composite score renormalizes.

MEWS scorecard: five vitals, one composite

MEWS, the composite. Hospitals learned that patients rarely crash one vital sign at a time; they deteriorate a little on several at once, and any single reading looks dismissable. The Modified Early Warning Score fixes that by scoring each vital 0-3 against its baseline and summing. NerfWatch applies the same construction to the five signals above: 0-3 points each, normalized to a percentage, WARNING at 50%, CRITICAL at 75%. A model with thinking down two points, inflation up two, and connectors down two sits at 6/15. That is 40%: no alarm yet, but a line visibly climbing on the dashboard, the early warning a threshold system cannot give you.

EWMA, the control chart. An exponentially weighted moving average holds a smoothed baseline for each signal, with control bands around it. An observation landing 2 sigma outside the band is a WARNING; 3 sigma is CRITICAL. This is the cliff detector: a 73% thinking-token drop blows through the 3-sigma band within one to three observations. The bands start wide and tighten as calibration accrues, which is why NerfWatch is silent for its first ten runs.

CUSUM, the drift detector. EWMA has a known blind spot: its baseline adapts, so a model losing 2% a week just drags the baseline down with it and never trips the band. The boiled frog. CUSUM accumulates every small deviation from the established level, and when the running total crosses a decision boundary (4 sigma, with 0.5 sigma of slack so noise doesn't accumulate), it fires a DRIFT alert. The cliff detector and the slope detector cover each other's blind spots.

BOCPD, the moment finder. EWMA and CUSUM tell you that something changed. Bayesian online changepoint detection tells you when: at each observation it computes the probability that a new regime just started, and records the changepoint with a confidence score. This is the instrument that turns "the model feels worse lately" into "behavior changed on Tuesday at 14:02, confidence 0.94." One is a hunch; the other is an incident report.

Probe scores. Layer 2's ground truth. Each probe has a known-correct answer computed in Go, graded deterministically (exact match, numeric tolerance, regex) plus a cross-provider LLM judge scoring the reasoning itself on a 0-1 rubric. Scores chart per dimension: logic, math, factual, and code separately, so a collapse in one cannot average out against the others.

Quick Start

# 1. Initialize - creates ~/.nerfwatch/, encrypted credstore, and config.
#    Walks you through enrolling providers interactively.
nerfwatch init

# 2. Run a single probe cycle to verify setup (reads keys from the credstore).
nerfwatch run

# 3. Start continuous monitoring with the heartbeat goroutine,
#    severity hysteresis, and the web dashboard.
nerfwatch monitor

Need to enroll more providers later? Use nerfwatch auth login <provider>. Suspect a model is regressing and want Gold-tier validation for a bounded window? Use nerfwatch burst --model <id> --cycles 5.

What you'll need

  • Go 1.25+ to build from source. No pre-built binaries for v0.1.
  • An API key for at least one supported provider (Anthropic, OpenAI, Google, or xAI). Cross-provider grading, and the full Gold tier, wants at least three enrolled providers.
  • macOS or Linux. Most commands run on Windows, but single-instance enforcement on nerfwatch monitor relies on POSIX flock and is a known gap on Windows.
  • Docker (optional). Required only if you run code probes, which execute model-generated code inside a restricted container.

How It Works

NerfWatch measurement architecture

Three independent detection layers, three statistical engines, one binary.

Layer 1: behavioral intrinsics. No test questions needed. NerfWatch reads the model's vitals off every response it makes: thinking-token depth, output inflation, latency shape, self-interruption frequency, and connector density in the reasoning text. Free signal, extracted from traffic you were already paying for.

Layer 2: probe-based quality scoring. Parameterized probes across four reasoning dimensions (logic, math, factual, code), graded for answer correctness and for chain-of-thought quality by a cross-provider judge. Code probes execute the model's generated code in a Docker sandbox against real test cases.

Layer 3: frozen anchors. A baseline captured at a moment you decided the model was good. Re-run the same probes weeks later and diff. Rolling baselines absorb slow decay; a frozen one remembers what good looked like. Alerts when the average delta exceeds a threshold (10% by default).

The engines. EWMA control charts for sudden shifts, CUSUM for sustained drift, BOCPD for pinpointing the change. All direction-aware, all silent through the ten-run calibration phase. Full definitions in What the numbers mean.

One degradation, three instruments: EWMA, CUSUM, BOCPD

Alerts. JSONL and SQLite always; webhooks, Slack, and email if configured. Per-dimension cooldowns keep one degradation event from becoming fifty notifications.

Cross-provider grading. No model grades itself, enforced in code. Four providers supported (Anthropic, OpenAI, Google, xAI); a degraded model cannot mask its own decline.

No model grades itself: four-provider cross-grading matrix

Limitations

NerfWatch is v0.1. Known gaps, all deliberate:

  • No usage telemetry, by design. NerfWatch does not emit or collect any usage data. The rationale and revisit triggers live in docs/brainstorms/nerfwatch-telemetry-requirements.md. If you adopt NerfWatch, please add yourself to ADOPTERS.md; self-reported adoption is the project's only adoption signal.
  • No pre-built binaries for v0.1. Install is git clone && go build. Cross-platform release binaries are post-launch work.
  • No SDK integration mode. NerfWatch monitors via standalone probes, not by wrapping production traffic. Planned for a future release.
  • No hosted dashboard. The dashboard is a local binary. A hosted multi-tenant version is explicitly post-0.1.
  • OpenAI contributes four behavioral signals, not five. OpenAI redacts reasoning text, so connector density is unavailable. The MEWS composite normalizes accordingly, but cross-provider comparison benefits from at least one full-reasoning provider (Anthropic, Google, or xAI).
  • Windows single-instance enforcement is not implemented. syscall.Flock is POSIX-only; running two nerfwatch monitor processes on Windows is not prevented.

Common questions

Phrased the way people actually ask them.

How can I tell if my model is getting dumber?

Not from vibes, and that is the trap: the output still reads fluent while reasoning depth quietly drops. The reliable tell is measurement over time. Ask the model the same parameterized questions on a schedule, score the answers, and chart reasoning effort (thinking tokens, connector density) against a baseline. That is what NerfWatch automates: if the model got dumber, the control charts show when it happened and on which dimension.

How do I detect AI model degradation or quality regression?

Three measurements that cover each other's blind spots: behavioral vital signs extracted free from every response, scheduled probes with known-correct answers for ground truth, and frozen baselines re-run weeks later for long-horizon drift. Statistical process control charts (EWMA, CUSUM, BOCPD) then decide whether a change is real or noise, so you are not eyeballing a dashboard and guessing.

Did the provider silently change the model behind my API?

It happens: retraining, quantization, routing changes, safety-layer updates, rarely announced. Changepoint detection (BOCPD) pinpoints the moment behavior shifted, with a confidence score. That turns "it feels different since last week" into "behavior changed on the 14th at 09:40, here is the chart," which is a very different conversation to have with a provider.

Why did my prompts suddenly stop working?

Either your workload changed or the model did, and without a monitor you cannot tell which. Daily probes give you the counterfactual: if the probes degraded too, it is the model, and you have the chart to prove it. If the probes held steady, it is your prompts, and you just saved days of blaming the wrong suspect.

Is my model being nerfed? Is it getting lazy?

The community has its own words for this: nerfed, lobotomized, lazy. The phenomenon is real and crosses vendors: GPT-4's "laziness" wave of December 2023 was real enough that OpenAI acknowledged and fixed it, and the Opus incident of early 2026 (a 73% thinking-token drop, unacknowledged for weeks) settled the question again. Whether any given week's "it got lazy" is real degradation or recency bias is the question a baseline answers. Receipts beat vibes.

Does this replace evals or LLM observability platforms?

No, and it does not try to. Eval suites measure how good a model is; this measures whether it changed. Observability platforms watch your production traffic from inside your application; this watches the model from outside, with zero code changes. They compose: many teams want all three.

Help and feedback

Installation

# Clone and build
git clone https://github.com/tendaigomo/NerfWatch.git
cd NerfWatch
go build -o nerfwatch .

# Move to your PATH (optional)
sudo mv nerfwatch /usr/local/bin/

Requires Go 1.25+ and an API key for at least one supported provider (Anthropic, OpenAI, Google, or xAI).

Commands

nerfwatch init

Declarative-first setup. Creates ~/.nerfwatch/, writes an encrypted credential store, generates nerfwatch.yaml from a manifest or an interactive TUI, and copies example probes. Interactive mode uses a Huh-based TUI so the triage room setup feels like hanging the chart above the bed, not filling out forms.

# Interactive (default) - walks you through providers, models, and a grader pairing.
nerfwatch init

# Declarative - reproducible installs from a manifest file.
nerfwatch init --manifest ./install.yaml --non-interactive

# CI-friendly - pipe the passphrase in; never echo it to stdout.
nerfwatch init --manifest ./install.yaml --non-interactive --passphrase-stdin < /run/secrets/nerfwatch

The wizard (or manifest) asks for:

  • Provider (anthropic, openai, google, or xai)
  • Model ID (e.g., claude-fable-5, gpt-5.6-sol, gemini-3.1-pro-preview, grok-4.5)
  • API key (entered securely, not echoed)
  • Grader model for CoT quality scoring (cross-provider recommended)

A single file-lock on ~/.nerfwatch/nerfwatch.lock ensures only one init runs at a time - re-running init on an existing install is safe and non-destructive. See docs/migration-v1-to-v2.md for the manifest schema and the v1 → v2 upgrade path.

API keys never land in nerfwatch.yaml. They are written to a passphrase-encrypted credstore at ~/.nerfwatch/credentials.enc (AES-256-GCM, Argon2id KDF). The YAML config stores only api_key_ref: pointers; see docs/auth.md for the credential lifecycle.

nerfwatch run

Executes a single probe cycle against all configured models. Useful for testing your setup and inspecting individual results.

nerfwatch run
nerfwatch run --config /path/to/nerfwatch.yaml

Prints a results table per model showing each probe's dimension, correctness score, and CoT quality score, plus a cost summary based on configured per-model token prices ($0 when unset).

Note: run stores results but does not update EWMA baselines or evaluate alerts. Use monitor for continuous statistical tracking.

nerfwatch monitor

Starts the continuous monitoring loop. This is the primary operating mode.

nerfwatch monitor

Each cycle:

  1. Executes all probes against all configured models
  2. Updates EWMA baselines per signal and dimension
  3. Evaluates alert thresholds
  4. Writes alerts to alerts.jsonl and SQLite

The first 10 runs are a calibration phase - NerfWatch establishes baselines before activating alerting. During calibration, the output shows progress (CALIBRATING 4/10 runs).

Only one monitor instance can run at a time (enforced by file lock on ~/.nerfwatch/nerfwatch.lock). Graceful shutdown on SIGINT/SIGTERM - completes the current probe before exiting.

nerfwatch dashboard

Web dashboard with time-series charts, alert investigation, and model comparison.

nerfwatch dashboard              # Web dashboard (default, opens browser)
nerfwatch dashboard --port 9000  # Custom port
nerfwatch dashboard --tui        # Terminal UI (original bubbletea dashboard)

The web dashboard is a React SPA embedded in the Go binary via embed.FS - no Node.js runtime required. It provides five views:

  • Overview - Health grid with MEWS scores, status badges, and 7-day sparklines per model
  • Model Detail - Time-series charts with EWMA control bands, CUSUM drift status, and BOCPD changepoint markers
  • Alerts - Filterable alert feed with drill-down to triggering probe runs
  • Probe Explorer - Paginated history of probe runs with grading breakdowns
  • Comparison - Side-by-side model health across all dimensions

Auto-refreshes via polling (30s) while nerfwatch monitor is running.

The terminal UI remains available via --tui for headless/SSH workflows. Keyboard: j/k to scroll, r to refresh, ? for help, q to quit.

nerfwatch status

One-line-per-model health summary, designed for scripting.

nerfwatch status
nerfwatch status --json

Output format:

claude-fable-5     NORMAL   MEWS:2/15   DIMS:ok/ok/ok       LAST_RUN:5m ago
gpt-5.6-sol        WARNING  MEWS:8/12   DIMS:ok/warn/ok     LAST_RUN:12m ago

Exit codes: 0 if all models are NORMAL, 1 if any WARNING, 2 if any CRITICAL.

nerfwatch alerts

Query alert history from the SQLite database.

nerfwatch alerts                          # last 10 alerts
nerfwatch alerts --last 20                # last 20
nerfwatch alerts --severity CRITICAL      # only critical
nerfwatch alerts --model claude-fable-5   # filter by model
nerfwatch alerts --json                   # structured output

nerfwatch auth

Manage the encrypted credential store. Nine subcommands cover the full lifecycle - enrolling a provider, listing what's stored, rotating keys, exporting the passphrase for backup, and migrating a v1 plaintext YAML into the v2 store.

nerfwatch auth login anthropic               # Enroll a provider (prompts for key)
nerfwatch auth list                          # Which providers are currently enrolled
nerfwatch auth verify                        # Ping each provider to prove the stored key still works
nerfwatch auth rotate anthropic              # Replace a key without restarting the monitor
nerfwatch auth remove anthropic              # Tombstone a credential
nerfwatch auth revoke anthropic              # Zero a credential and signal the running monitor
nerfwatch auth reset --confirm               # Nuke the credstore (requires explicit confirmation)
nerfwatch auth migrate                       # One-shot import from v1 plaintext nerfwatch.yaml
nerfwatch auth export-passphrase             # Print the passphrase for backup (mode 0600 recommended)

Every subcommand routes through the atomic KeyProvider - a write flushes the credstore to disk, syncs the parent directory, and only then returns success. See docs/auth.md for threat model, rotation semantics, and the detailed failure-mode matrix.

nerfwatch burst

Pin a single model to a higher grading tier for a bounded window of cycles. Typical use: a suspected regression on a Silver-tier model warrants Gold-tier validation for five cycles before deciding on further action. Think of it as ordering a stat echocardiogram when the routine monitor flags an anomaly - more evidence, same patient, no paging the full cardiac team.

nerfwatch burst --model claude-fable-5 --cycles 5 --tier gold
nerfwatch burst --cancel                                  # Abort an active burst, no monitor restart

burst writes ~/.nerfwatch/burst.json (mode 0600) signed with a per-install HMAC token at ~/.nerfwatch/.burst_token. The monitor verifies the signature every cycle and ignores a tampered file with a warn-level log. The web dashboard shows a BURST badge on the affected model card and a banner with the remaining cycles - cancel is idempotent and safe to run if you're not sure whether a burst is active.

nerfwatch configure add-model

Add a new model to an existing configuration.

nerfwatch configure add-model

Interactive prompt for provider, model ID, API key, and grader configuration. Validates that the grader model differs from the model under test.

nerfwatch purge

Remove old monitoring data from the SQLite database.

nerfwatch purge --older-than 90d              # dry run (default)
nerfwatch purge --older-than 90d --confirm    # actually delete

Protects baseline data - refuses to purge records that overlap with active baseline computations.

nerfwatch sandbox build

Build the Docker image used by code probes. Required before code probes can execute.

nerfwatch sandbox build

The image includes Python 3, Node.js (JavaScript/TypeScript via tsx), JDK (Java), g++ (C++), and Go. Requires Docker to be installed and running.

nerfwatch sandbox status

Check sandbox readiness - whether Docker is available and the sandbox image exists.

nerfwatch sandbox status

nerfwatch seed

Generate realistic demo data for the web dashboard. Creates 30 days of multi-model time-series data including regression events, gradual drift, and corresponding alerts.

nerfwatch seed              # Generate demo data
nerfwatch seed --force      # Regenerate (overwrites existing demo data)

Uses synthetic model IDs with a demo- prefix (demo-claude-opus, demo-gpt-4o, demo-gemini-pro) that never collide with real monitoring data.

nerfwatch anchor

Manage Layer 3 frozen baselines for long-term drift detection.

nerfwatch anchor capture --model claude-fable-5 # Freeze current scores as baseline
nerfwatch anchor check --model claude-fable-5   # Re-run probes, compare to frozen baseline
nerfwatch anchor status                         # Display all anchor status and results

nerfwatch changepoints

View BOCPD-detected regime changes.

nerfwatch changepoints                          # Recent changepoints
nerfwatch changepoints --last 20                # Last 20
nerfwatch changepoints --model claude-fable-5   # Filter by model
nerfwatch changepoints --signal thinking_tokens # Filter by signal

Configuration

NerfWatch uses a hierarchical configuration system with three layers (highest priority first):

  1. Environment variables - nested overrides like NERFWATCH_DB_PATH, NERFWATCH_SCHEDULE_INTERVAL, provider key overrides like NERFWATCH_ANTHROPIC_API_KEY, and provider endpoint overrides like NERFWATCH_ANTHROPIC_BASE_URL
  2. YAML config file - ~/.nerfwatch/nerfwatch.yaml (or --config flag)
  3. Compiled defaults - sensible defaults for all settings

Routing through a local proxy (optional)

Each model (and grader) accepts an optional base_url, and each provider accepts a NERFWATCH_<PROVIDER>_BASE_URL env override (env wins). This routes probes through any OpenAI-/Anthropic-compatible endpoint - for example a local CLIProxyAPI sidecar that re-exposes a flat-rate Claude or Codex subscription:

models:
  - name: gpt-5.6-sol
    provider: openai
    model_id: gpt-5.6-sol
    base_url: "http://localhost:8317"   # host only - NerfWatch appends /v1/... itself
    api_key_ref: "credentials://openai"

Two rules, both enforced at config load:

  • base_url must be an absolute http(s) URL, host only - a trailing /v1 or /v1beta is rejected because the adapters append the version path themselves (a suffix would yield a doubled path and silent 404s).
  • Keep any local proxy bound to loopback; it is a gateway to a paid account.

Before trusting proxied samples in baselines, run the usage-parity smoke test (TestSidecarUsageParity in internal/provider/) - if a proxy strips token usage, its samples must never enter an SPC baseline. nerfwatch auth verify always targets the provider's public endpoint regardless of base_url.

Credentials are never stored in the YAML file. They live in a passphrase-encrypted credstore at ~/.nerfwatch/credentials.enc (AES-256-GCM, Argon2id KDF) and the YAML references them by api_key_ref: pointers. This is the v2 format; legacy v1 configs with inline api_key: fields are migrated automatically on first nerfwatch init or explicitly via nerfwatch auth migrate. See docs/migration-v1-to-v2.md.

Declarative setup

For reproducible installs, nerfwatch init --manifest consumes a small YAML manifest that lists the providers to enroll and the grader pairings to write. The manifest never contains secrets; keys are either prompted for interactively, piped in via --passphrase-stdin, or read from pre-populated environment variables (NERFWATCH_ANTHROPIC_API_KEY, etc.). See the manifest schema in docs/migration-v1-to-v2.md.

Example configuration

# v2 format: api_key_ref points into the encrypted credstore.
# Legacy v1 files with inline api_key: are auto-migrated on `nerfwatch init`.
models:
  - name: claude-fable-5
    provider: anthropic
    model_id: claude-fable-5
    api_key_ref: "credentials://anthropic"
    schedule_interval: "30m"          # Per-model override (expensive, run less often)
    grader:
      provider: openai
      model_id: gpt-5.4-mini
      api_key_ref: "credentials://openai"

  - name: gpt-5.6-sol
    provider: openai
    model_id: gpt-5.6-sol
    api_key_ref: "credentials://openai"
    grader:
      provider: anthropic
      model_id: claude-haiku-4-5-20251001
      api_key_ref: "credentials://anthropic"

  - name: gemini-3.1-pro
    provider: google
    model_id: gemini-3.1-pro-preview
    api_key_ref: "credentials://google"
    grader:
      provider: openai
      model_id: gpt-5.4-mini
      api_key_ref: "credentials://openai"

  - name: grok-4.5
    provider: xai
    model_id: grok-4.5
    api_key_ref: "credentials://xai"
    grader:
      provider: anthropic
      model_id: claude-haiku-4-5-20251001
      api_key_ref: "credentials://anthropic"

probe_file: probes.yaml
db_path: ~/.nerfwatch/nerfwatch.db

schedule:
  interval: "15m"                     # Global default

alerts:
  cooldown_hours: 4
  jsonl_path: ~/.nerfwatch/alerts.jsonl
  channels:                           # External alert delivery (optional)
    - type: webhook
      url: "https://example.com/nerfwatch-alerts"
    - type: slack
      url: "https://hooks.slack.com/services/T.../B.../..."
    - type: email
      smtp_host: smtp.gmail.com
      smtp_port: 587
      smtp_from: nerfwatch@example.com
      smtp_to: ops@example.com
      smtp_user: nerfwatch@example.com
      smtp_pass: "app-password"

anchor:
  check_interval: "168h"              # Weekly re-check
  alert_threshold: 0.10               # 10% delta triggers alert

Cross-provider grading

NerfWatch enforces that a model never grades its own output (a model that has degraded would rate its own degraded output as acceptable). Cross-provider grading is recommended: use an OpenAI model to grade Anthropic output and vice versa. Same-provider different-model grading (e.g., Haiku grading Opus) is permitted as a fallback.

Grading tiers

NerfWatch labels every grading cycle with a tier based on enrolled providers (N) and peers-per-probe:

Tier N peers/probe Monitor behavior
Gold ≥ 3 N-1 Every peer grades every target - strongest confidence.
Silver ≥ 4 2 Fixed two-peer grading - reduced cost, still cross-provider.
Bronze ≥ 3 1 One peer per probe - minimum viable cross-grading.
Deterministic-only < 2 - LLM-judge disabled; falls back to exact / regex / numeric graders. Dashboard shows a red banner.

When N drops below 3, the dashboard shows a cross-grading degraded banner so you know probe-quality confidence has fallen. Below N=2 the banner goes red and LLM-judge probes shut off entirely. Restore confidence by running nerfwatch auth login <provider> to enroll more. See docs/grading-tiers.md for the full tier math and promotion rules.

Custom probes

The example probes (probes.example.yaml) ship with known answers that models may have seen during training. For meaningful monitoring, create your own probes.yaml with domain-specific reasoning challenges.

Probe format:

version: "1.0.0"
probes:
  - id: "my-logic-probe-01"
    version: "1.0.0"
    dimension: "logic"          # logic, math, factual, or code
    prompt: "Given {{.premise}}, what follows about {{.subject}}?"
    variables:
      premise: ["All X are Y", "No X are Y", "Some X are Y"]
      subject: ["item A", "item B", "item C"]
    answer_func: "logic-syllogism"  # registered Go function
    grading_method: "llm_judge"     # exact, numeric_tolerance, regex, or llm_judge

Each run selects random values from the variable pools using a stored seed, so probes are reproducible but vary across runs to prevent memorization.

Architecture

nerfwatch/
  main.go                          # Entry point
  cmd/                             # Cobra CLI commands
    root.go                        #   Root command + global flags
    init.go                        #   nerfwatch init
    run.go                         #   nerfwatch run
    monitor.go                     #   nerfwatch monitor (main loop)
    dashboard.go                   #   nerfwatch dashboard (web + TUI)
    status.go                      #   nerfwatch status (one-liner)
    alerts.go                      #   nerfwatch alerts (history)
    configure.go                   #   nerfwatch configure
    configure_addmodel.go          #   nerfwatch configure add-model
    purge.go                       #   nerfwatch purge
    sandbox.go                     #   nerfwatch sandbox (build/status)
    seed.go                        #   nerfwatch seed (demo data generation)
    anchor.go                      #   nerfwatch anchor (capture/check/status)
    changepoints.go                #   nerfwatch changepoints (BOCPD results)
  internal/
    config/                        # Hierarchical config (koanf)
    storage/                       # SQLite with WAL mode, dual pools
    provider/                      # Provider interface + adapters
      provider.go                  #   Request/Response types, Provider interface
      anthropic.go                 #   Anthropic adapter (ThinkingBlock extraction)
      openai.go                    #   OpenAI adapter (ReasoningTokens)
      google.go                    #   Google Gemini adapter (thinking parts)
      xai.go                       #   xAI Grok adapter (reasoning_content)
      registry.go                  #   Provider name -> implementation map
    probe/                         # Probe system
      template.go                  #   ProbeTemplate, GeneratedProbe, Generate()
      loader.go                    #   YAML loader with strict validation
      answer_funcs.go              #   Answer function registry (8 MVP functions)
      answer_funcs_code.go         #   Code probe answer functions (5 probes)
      engine.go                    #   Execution pipeline (sequential + concurrent grading)
    grading/                       # Grading system
      deterministic.go             #   Exact, numeric tolerance, regex graders
      llm_judge.go                 #   LLM-as-judge with rubric + XML delimiters
      sandbox.go                   #   Docker sandbox grader for code probes
    sandbox/                       # Docker sandbox for code execution
      sandbox.go                   #   Container lifecycle (create, exec, teardown)
      testspec.go                  #   TestSpec type and JSON serialization
      extractor.go                 #   Code extraction from LLM responses
      harness.go                   #   Embedded per-language test runners
    observe/                       # Behavioral signal extraction
      intrinsics.go                #   5 Layer 1 signals from API responses
      mews.go                      #   MEWS composite score computation
    stats/                         # Statistical engines
      ewma.go                      #   EWMA tracker with time-varying control limits
      cusum.go                     #   CUSUM two-sided drift detection
      bocpd.go                     #   Bayesian Online Changepoint Detection
      baseline.go                  #   Per-model per-signal baseline management
    anchor/                        # Layer 3 frozen baselines
      anchor.go                    #   Capture, check, and status operations
    alert/                         # Alert system
      alert.go                     #   Evaluator with per-dimension cooldown
      writer.go                    #   JSONL file writer (mode 0600)
      channels.go                  #   Webhook, Slack, email delivery
    server/                        # Web dashboard backend
      server.go                    #   HTTP server with embedded SPA
      api.go                       #   REST API handlers (9 endpoints)
    dashboard/                     # Terminal UI
      model.go                     #   bubbletea Model (Init/Update/View)
      views.go                     #   lipgloss rendering + status formatting
  web/                             # React SPA (compiled, embedded via embed.FS)
  probes.example.yaml              # 13 example probes (3 logic, 3 math, 2 factual, 5 code)
  nerfwatch.example.yaml           # Example configuration

Data flow

nerfwatch monitor
  |
  v
Probe Engine ──> Provider Adapter ──> AI Model API
  |                    |
  |              Response + Usage
  |                    |
  +──> Behavioral Intrinsics Observer ──> Layer 1 signals
  |         |
  |         +──> MEWS Composite Score
  |
  +──> Grading System ──> Correctness + CoT quality
  |         |
  |         +──> Docker Sandbox (code probes)
  |
  +──> SQLite Storage (probe_runs, probe_results, intrinsic_observations)
  |
  v
Baseline Manager ──> EWMA Update ──> Threshold Check
  |                  CUSUM Update ──> Drift Detection
  |                  BOCPD ────────> Changepoint Detection
  v
Alert Evaluator ──> Per-dimension cooldown check
  |
  +──> Internal: alerts.jsonl + SQLite alerts table
  +──> External: Webhook + Slack + Email (configured channels)
  |
  v
Web Dashboard <── REST API (9 endpoints) <── SQLite read pool

Storage

NerfWatch uses SQLite with WAL mode for concurrent read/write access. The database lives at ~/.nerfwatch/nerfwatch.db by default.

  • Write pool: single connection (MaxOpenConns=1) for serialized writes
  • Read pool: 10 connections for dashboard, status, and alert queries
  • Transactions: all writes use BEGIN IMMEDIATE for fast failure on contention
  • Schema: 9 tables (probe_runs, probe_results, intrinsic_observations, baselines, cusum_trackers, alerts, changepoints, anchors, anchor_results) with indices

Statistical engines

NerfWatch uses three complementary detection engines from Statistical Process Control:

EWMA (Exponentially Weighted Moving Average) - catches sudden shifts:

  • Lambda: 0.25 for Layer 2 probe dimensions, 0.2 for behavioral intrinsics
  • Control limits: exact time-varying limits (wider during warm-up, converge to steady-state)
  • Thresholds: WARNING at 2-sigma breach, CRITICAL at 3-sigma breach
  • Calibration: 10 runs within a 7-day window required before alerting activates

CUSUM (Cumulative Sum) - catches sustained drift:

  • K: 0.5 sigma (slack parameter)
  • H: 4.0 sigma (decision boundary)
  • Two-sided: independent upper and lower accumulators
  • Fires DRIFT alerts when accumulated deviation exceeds H, independent of EWMA

BOCPD (Bayesian Online Changepoint Detection) - pinpoints regime changes:

  • Identifies the exact moment a signal's behavior changes
  • Records changepoints with confidence scores and affected dimensions
  • Queryable via nerfwatch changepoints

All three engines are direction-aware: some signals alert on decrease (thinking tokens), others on increase (inflation ratio), others bidirectionally (latency). A 73% drop in a signal (like the Opus regression) triggers EWMA CRITICAL within 1-3 observations and CUSUM DRIFT shortly after.

Behavioral signals

Signal What it measures Direction Provider support
ThinkingTokens Chain-of-thought depth (word count proxy) Below only Anthropic (full), OpenAI (count only), Google (full), xAI (full)
InflationRatio Output tokens / input tokens Above only All
LatencyProfile Milliseconds per output token Bidirectional All
EarlyStopping Self-interruption phrases ("Can I continue?") Above only All
ConnectorDensity Logical connectors per 100 words in thinking Below only Anthropic, Google, xAI (not OpenAI)

OpenAI models contribute 4 signals (no ConnectorDensity since reasoning content isn't exposed). Google and xAI provide full thinking/reasoning text, enabling all 5 signals like Anthropic. MEWS normalizes accordingly.

How NerfWatch Compares

NerfWatch LangSmith Patronus AI
Approach Standalone probes - zero code changes SDK integration + tracing SDK integration + guardrails
Detection Statistical Process Control (EWMA + CUSUM + BOCPD) Threshold-based evals Rule-based + model evals
Catches Unknown degradation you didn't anticipate Known failures you define Policy violations you specify
Metaphor ICU cardiac monitor (always watching) Lab blood test (you schedule it) Security guard (checks rules)
Setup nerfwatch init && nerfwatch monitor Instrument every LLM call Instrument every LLM call
Cost API keys only Platform subscription Platform subscription

Security

  • Config file (nerfwatch.yaml) and database created with mode 0600
  • Data directory (~/.nerfwatch/) created with mode 0700
  • Alerts file (alerts.jsonl) created with mode 0600, contains derived metrics only - never raw model output
  • API keys are encrypted at rest in the credstore (~/.nerfwatch/credentials.enc, AES-256-GCM + Argon2id); the YAML config stores only opaque api_key_ref: pointers. Env var override available (NERFWATCH_*_API_KEY).
  • Burst-window state file (~/.nerfwatch/burst.json) is HMAC-signed with a per-install token so a tampered file is ignored by the monitor (warn-level log, no crash).
  • LLM judge input wrapped in XML delimiters to mitigate prompt injection from model-under-test output

Roadmap

v0.1 - public launch (next)

Current main is feature-complete for a first public release: init + encrypted credentials, cross-grading engine, monitor with heartbeat and severity hysteresis, calibration report, web dashboard, seed command, full CI matrix (Go + Vitest + Playwright across chromium/firefox/webkit/mobile/visual/lighthouse). Telemetry deliberately declined for now - see docs/brainstorms/nerfwatch-telemetry-requirements.md.

What's left before tagging v0.1 is launch prep, not product work: README polish for a cold-reading first adopter, docs review pass, decision on tendaigomo/NerfWatch visibility flip, release tag + notes, and an adopter outreach plan.

Post-0.1 (unplanned order)

  • SDK wrapper integration mode (monitor production traffic without separate probes).
  • Public hosted dashboard at nerfwatch.dev.
  • User authentication and multi-tenant data isolation.
  • WebSocket push (replace polling).
  • Dashboard-driven configuration editing.
  • Crash telemetry (Sentry-style error reporting - separate decision from the usage-telemetry rejection, needs its own brainstorm).

Contributing

We welcome contributions! See CONTRIBUTING.md for guidelines on adding providers, probes, and statistical engines.

License

MIT - see LICENSE for details.

About

The Taravangian Test for AI: catch silent model degradation before your users do. SPC-based reasoning-quality monitoring for Claude, GPT, Gemini, and Grok.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages