The Taravangian Test for AI.
Catch silent model degradation before your users do.
In Brandon Sanderson's Stormlight Archive, King Taravangian wakes each day not knowing whether he'll be a genius or barely able to function. So every morning he takes the same test, measures the result, and only then decides what he can be trusted to do that day.
AI models have the same condition. Providers retrain, requantize, and reroute without telling you, so the model you built against is not guaranteed to be the model serving your users right now. The only way to know which one woke up is to test it. Same test, every morning, measured.
NerfWatch is that test, run as a monitor. It probes your models on a schedule, pulls behavioral vital signs out of every response, and feeds the numbers through the same statistical process control charts hospitals use to watch patients. When a chart fires, you get an alert that names the signal, the severity, and the moment things changed.
A standalone Go binary that watches AI models from the outside. No SDK in your code, no proxy in your request path. You give it API keys; it sends parameterized probes on its own schedule, grades every answer with a different provider's model, and charts five behavioral signals over time: thinking tokens, output inflation, latency, early stopping, connector density.
This is not a one-vendor problem. OpenAI users spent December 2023 documenting GPT-4's "laziness" until OpenAI acknowledged it and shipped a fix. In April 2025 an update made GPT-4o sycophantic enough that OpenAI rolled it back and published a postmortem. And in early 2026, Claude Opus silently lost most of its reasoning depth: thinking tokens fell 73%, reads-per-edit fell from 6.6 to 2.0, and the proof came not from the provider but from one engineer who hand-analyzed 6,852 sessions. Three incidents, two vendors, one pattern. The next one can come from any of them, which is why NerfWatch monitors Anthropic, OpenAI, Google, and xAI models with identical instruments and no favorites. It exists so the next proof takes one to three probe cycles, not six thousand sessions.
The frontier has moved on since those incidents, and visibility has gotten worse, not better. Fable 5 never returns its raw chain of thought; you get a summary or nothing. GPT-5.6 Sol reports reasoning as a bare token count. Gemini 3.1 Pro and Grok 4.5 still expose reasoning text today, with no guarantee about tomorrow. The stronger the model, the less of its thinking you are allowed to see. What's left is the outside view: measure behavior the way a doctor reads vitals off a patient who can't tell you how they feel. That is the whole product, and it works the same on all four providers.
These follow from the Taravangian condition. You can't know which model woke up, so the design assumes you have to find out every cycle:
- Test every cycle, not once. A model that passed your evaluation at launch can quietly become a different model tomorrow. NerfWatch re-takes the test on a schedule and never assumes yesterday's result still holds.
- A degraded model can't grade itself. The bent ruler cannot measure its own bend, so grading is always cross-provider. No model scores its own output.
- Measure each dimension, not an average. The Opus collapse was about 4% as a single aggregate score, dismissable noise, yet a 25-point drop in chain-of-thought depth. NerfWatch tracks logic, math, factual, and code separately so a collapse in one cannot hide behind health in the others.
- Statistics, not thresholds. EWMA, CUSUM, and BOCPD catch the degradation you didn't anticipate, not just the failures you remembered to write a rule for.
- Private, parameterized probes. Any published test set is eventually trained on. NerfWatch's probes vary per run from a stored seed: reproducible, never memorizable.
- Derived metrics only, local-first. NerfWatch records statistical signals, never raw model output, in a SQLite database on your machine. No cloud, no telemetry.
| Scenario | Layer | Detection | Time to alert |
|---|---|---|---|
| Opus-style reasoning collapse (73% thinking-token drop) | Layer 1 MEWS | CRITICAL | 1-3 runs |
| Gradual CoT quality erosion (10% over 2 weeks) | Layer 2 + CUSUM | DRIFT | Days, not weeks |
| Provider hides thinking tokens | Layer 1 signal loss | WARNING | Immediate |
| Code generation correctness regression | Layer 2 code probes | WARNING/CRITICAL | 1-3 runs |
| Model upgrade changes behavior | Layer 1 + BOCPD | Changepoint | Exact moment identified |
Every NerfWatch alert names the instrument that fired. None of these are invented for this project; they come from clinical wards and factory floors, where the cost of a missed shift is measured in lives and recalls. Here is each one, what it measures, and why it earns its place.
Thinking tokens. How much reasoning the model did before answering. Counted from reasoning text where the provider exposes it, from usage metadata where it doesn't. This was the first visible symptom of the Opus collapse: a model that used to spend 1,200 tokens working a logic probe started spending 300. Whatever the final answer says, that is a different model. Alerts fire on drops; a model thinking more is not an emergency.
Inflation ratio. Output tokens relative to input. Degrading models often get wordier while getting shallower; verbosity fills the space reasoning used to occupy. During the Opus incident, output volume rose while thinking fell. Watching both catches the trade. Alerts fire on rises.
Latency profile. Milliseconds per output token. A sudden change here usually means the serving stack changed underneath you: different hardware, different quantization, different routing. That is not degradation by itself, which is why latency votes in the composite score instead of raising alarms alone.
Early stopping. How often the model interrupts its own work ("Shall I continue?", answers that trail off). A model that increasingly asks permission instead of finishing is failing in a way accuracy tests miss, because the answers it does complete may still be correct.
Connector density. Logical connectors (therefore, because, however) per 100 words of reasoning text. Sound reasoning is argument, and argument leaves scaffolding. When chain-of-thought decays from argument into narration, the connectors thin out before the answers go wrong, so this is often the earliest signal to move. Not measurable on OpenAI models, which don't expose reasoning text; the composite score renormalizes.
MEWS, the composite. Hospitals learned that patients rarely crash one vital sign at a time; they deteriorate a little on several at once, and any single reading looks dismissable. The Modified Early Warning Score fixes that by scoring each vital 0-3 against its baseline and summing. NerfWatch applies the same construction to the five signals above: 0-3 points each, normalized to a percentage, WARNING at 50%, CRITICAL at 75%. A model with thinking down two points, inflation up two, and connectors down two sits at 6/15. That is 40%: no alarm yet, but a line visibly climbing on the dashboard, the early warning a threshold system cannot give you.
EWMA, the control chart. An exponentially weighted moving average holds a smoothed baseline for each signal, with control bands around it. An observation landing 2 sigma outside the band is a WARNING; 3 sigma is CRITICAL. This is the cliff detector: a 73% thinking-token drop blows through the 3-sigma band within one to three observations. The bands start wide and tighten as calibration accrues, which is why NerfWatch is silent for its first ten runs.
CUSUM, the drift detector. EWMA has a known blind spot: its baseline adapts, so a model losing 2% a week just drags the baseline down with it and never trips the band. The boiled frog. CUSUM accumulates every small deviation from the established level, and when the running total crosses a decision boundary (4 sigma, with 0.5 sigma of slack so noise doesn't accumulate), it fires a DRIFT alert. The cliff detector and the slope detector cover each other's blind spots.
BOCPD, the moment finder. EWMA and CUSUM tell you that something changed. Bayesian online changepoint detection tells you when: at each observation it computes the probability that a new regime just started, and records the changepoint with a confidence score. This is the instrument that turns "the model feels worse lately" into "behavior changed on Tuesday at 14:02, confidence 0.94." One is a hunch; the other is an incident report.
Probe scores. Layer 2's ground truth. Each probe has a known-correct answer computed in Go, graded deterministically (exact match, numeric tolerance, regex) plus a cross-provider LLM judge scoring the reasoning itself on a 0-1 rubric. Scores chart per dimension: logic, math, factual, and code separately, so a collapse in one cannot average out against the others.
# 1. Initialize - creates ~/.nerfwatch/, encrypted credstore, and config.
# Walks you through enrolling providers interactively.
nerfwatch init
# 2. Run a single probe cycle to verify setup (reads keys from the credstore).
nerfwatch run
# 3. Start continuous monitoring with the heartbeat goroutine,
# severity hysteresis, and the web dashboard.
nerfwatch monitorNeed to enroll more providers later? Use nerfwatch auth login <provider>. Suspect a model is regressing and want Gold-tier validation for a bounded window? Use nerfwatch burst --model <id> --cycles 5.
- Go 1.25+ to build from source. No pre-built binaries for v0.1.
- An API key for at least one supported provider (Anthropic, OpenAI, Google, or xAI). Cross-provider grading, and the full Gold tier, wants at least three enrolled providers.
- macOS or Linux. Most commands run on Windows, but single-instance enforcement on
nerfwatch monitorrelies on POSIXflockand is a known gap on Windows. - Docker (optional). Required only if you run code probes, which execute model-generated code inside a restricted container.
Three independent detection layers, three statistical engines, one binary.
Layer 1: behavioral intrinsics. No test questions needed. NerfWatch reads the model's vitals off every response it makes: thinking-token depth, output inflation, latency shape, self-interruption frequency, and connector density in the reasoning text. Free signal, extracted from traffic you were already paying for.
Layer 2: probe-based quality scoring. Parameterized probes across four reasoning dimensions (logic, math, factual, code), graded for answer correctness and for chain-of-thought quality by a cross-provider judge. Code probes execute the model's generated code in a Docker sandbox against real test cases.
Layer 3: frozen anchors. A baseline captured at a moment you decided the model was good. Re-run the same probes weeks later and diff. Rolling baselines absorb slow decay; a frozen one remembers what good looked like. Alerts when the average delta exceeds a threshold (10% by default).
The engines. EWMA control charts for sudden shifts, CUSUM for sustained drift, BOCPD for pinpointing the change. All direction-aware, all silent through the ten-run calibration phase. Full definitions in What the numbers mean.
Alerts. JSONL and SQLite always; webhooks, Slack, and email if configured. Per-dimension cooldowns keep one degradation event from becoming fifty notifications.
Cross-provider grading. No model grades itself, enforced in code. Four providers supported (Anthropic, OpenAI, Google, xAI); a degraded model cannot mask its own decline.
NerfWatch is v0.1. Known gaps, all deliberate:
- No usage telemetry, by design. NerfWatch does not emit or collect any usage data. The rationale and revisit triggers live in docs/brainstorms/nerfwatch-telemetry-requirements.md. If you adopt NerfWatch, please add yourself to ADOPTERS.md; self-reported adoption is the project's only adoption signal.
- No pre-built binaries for v0.1. Install is
git clone && go build. Cross-platform release binaries are post-launch work. - No SDK integration mode. NerfWatch monitors via standalone probes, not by wrapping production traffic. Planned for a future release.
- No hosted dashboard. The dashboard is a local binary. A hosted multi-tenant version is explicitly post-0.1.
- OpenAI contributes four behavioral signals, not five. OpenAI redacts reasoning text, so connector density is unavailable. The MEWS composite normalizes accordingly, but cross-provider comparison benefits from at least one full-reasoning provider (Anthropic, Google, or xAI).
- Windows single-instance enforcement is not implemented.
syscall.Flockis POSIX-only; running twonerfwatch monitorprocesses on Windows is not prevented.
Phrased the way people actually ask them.
Not from vibes, and that is the trap: the output still reads fluent while reasoning depth quietly drops. The reliable tell is measurement over time. Ask the model the same parameterized questions on a schedule, score the answers, and chart reasoning effort (thinking tokens, connector density) against a baseline. That is what NerfWatch automates: if the model got dumber, the control charts show when it happened and on which dimension.
Three measurements that cover each other's blind spots: behavioral vital signs extracted free from every response, scheduled probes with known-correct answers for ground truth, and frozen baselines re-run weeks later for long-horizon drift. Statistical process control charts (EWMA, CUSUM, BOCPD) then decide whether a change is real or noise, so you are not eyeballing a dashboard and guessing.
It happens: retraining, quantization, routing changes, safety-layer updates, rarely announced. Changepoint detection (BOCPD) pinpoints the moment behavior shifted, with a confidence score. That turns "it feels different since last week" into "behavior changed on the 14th at 09:40, here is the chart," which is a very different conversation to have with a provider.
Either your workload changed or the model did, and without a monitor you cannot tell which. Daily probes give you the counterfactual: if the probes degraded too, it is the model, and you have the chart to prove it. If the probes held steady, it is your prompts, and you just saved days of blaming the wrong suspect.
The community has its own words for this: nerfed, lobotomized, lazy. The phenomenon is real and crosses vendors: GPT-4's "laziness" wave of December 2023 was real enough that OpenAI acknowledged and fixed it, and the Opus incident of early 2026 (a 73% thinking-token drop, unacknowledged for weeks) settled the question again. Whether any given week's "it got lazy" is real degradation or recency bias is the question a baseline answers. Receipts beat vibes.
No, and it does not try to. Eval suites measure how good a model is; this measures whether it changed. Observability platforms watch your production traffic from inside your application; this watches the model from outside, with zero code changes. They compose: many teams want all three.
- Bug reports and feature requests: https://github.com/tendaigomo/NerfWatch/issues
- Discussions (adoption friction, usage questions, design Q&A): https://github.com/tendaigomo/NerfWatch/discussions
- Using NerfWatch? Add yourself to ADOPTERS.md. It is the only signal feeding the telemetry-revisit decision.
- Release history: CHANGELOG.md
- Security policy: SECURITY.md. Report vulnerabilities through GitHub Security Advisories, not public issues.
- Contributing: CONTRIBUTING.md covers provider adapters, probe contributions, and statistical-engine work.
- Design decisions: DESIGN-DECISIONS/ is the judgment trail; docs/DESIGN_DECISIONS.md has the full engineering rationale.
# Clone and build
git clone https://github.com/tendaigomo/NerfWatch.git
cd NerfWatch
go build -o nerfwatch .
# Move to your PATH (optional)
sudo mv nerfwatch /usr/local/bin/Requires Go 1.25+ and an API key for at least one supported provider (Anthropic, OpenAI, Google, or xAI).
Declarative-first setup. Creates ~/.nerfwatch/, writes an encrypted credential store, generates nerfwatch.yaml from a manifest or an interactive TUI, and copies example probes. Interactive mode uses a Huh-based TUI so the triage room setup feels like hanging the chart above the bed, not filling out forms.
# Interactive (default) - walks you through providers, models, and a grader pairing.
nerfwatch init
# Declarative - reproducible installs from a manifest file.
nerfwatch init --manifest ./install.yaml --non-interactive
# CI-friendly - pipe the passphrase in; never echo it to stdout.
nerfwatch init --manifest ./install.yaml --non-interactive --passphrase-stdin < /run/secrets/nerfwatchThe wizard (or manifest) asks for:
- Provider (anthropic, openai, google, or xai)
- Model ID (e.g.,
claude-fable-5,gpt-5.6-sol,gemini-3.1-pro-preview,grok-4.5) - API key (entered securely, not echoed)
- Grader model for CoT quality scoring (cross-provider recommended)
A single file-lock on ~/.nerfwatch/nerfwatch.lock ensures only one init runs at a time - re-running init on an existing install is safe and non-destructive. See docs/migration-v1-to-v2.md for the manifest schema and the v1 → v2 upgrade path.
API keys never land in nerfwatch.yaml. They are written to a passphrase-encrypted credstore at ~/.nerfwatch/credentials.enc (AES-256-GCM, Argon2id KDF). The YAML config stores only api_key_ref: pointers; see docs/auth.md for the credential lifecycle.
Executes a single probe cycle against all configured models. Useful for testing your setup and inspecting individual results.
nerfwatch run
nerfwatch run --config /path/to/nerfwatch.yamlPrints a results table per model showing each probe's dimension, correctness score, and CoT quality score, plus a cost summary based on configured per-model token prices ($0 when unset).
Note: run stores results but does not update EWMA baselines or evaluate alerts. Use monitor for continuous statistical tracking.
Starts the continuous monitoring loop. This is the primary operating mode.
nerfwatch monitorEach cycle:
- Executes all probes against all configured models
- Updates EWMA baselines per signal and dimension
- Evaluates alert thresholds
- Writes alerts to
alerts.jsonland SQLite
The first 10 runs are a calibration phase - NerfWatch establishes baselines before activating alerting. During calibration, the output shows progress (CALIBRATING 4/10 runs).
Only one monitor instance can run at a time (enforced by file lock on ~/.nerfwatch/nerfwatch.lock). Graceful shutdown on SIGINT/SIGTERM - completes the current probe before exiting.
Web dashboard with time-series charts, alert investigation, and model comparison.
nerfwatch dashboard # Web dashboard (default, opens browser)
nerfwatch dashboard --port 9000 # Custom port
nerfwatch dashboard --tui # Terminal UI (original bubbletea dashboard)The web dashboard is a React SPA embedded in the Go binary via embed.FS - no Node.js runtime required. It provides five views:
- Overview - Health grid with MEWS scores, status badges, and 7-day sparklines per model
- Model Detail - Time-series charts with EWMA control bands, CUSUM drift status, and BOCPD changepoint markers
- Alerts - Filterable alert feed with drill-down to triggering probe runs
- Probe Explorer - Paginated history of probe runs with grading breakdowns
- Comparison - Side-by-side model health across all dimensions
Auto-refreshes via polling (30s) while nerfwatch monitor is running.
The terminal UI remains available via --tui for headless/SSH workflows. Keyboard: j/k to scroll, r to refresh, ? for help, q to quit.
One-line-per-model health summary, designed for scripting.
nerfwatch status
nerfwatch status --jsonOutput format:
claude-fable-5 NORMAL MEWS:2/15 DIMS:ok/ok/ok LAST_RUN:5m ago
gpt-5.6-sol WARNING MEWS:8/12 DIMS:ok/warn/ok LAST_RUN:12m ago
Exit codes: 0 if all models are NORMAL, 1 if any WARNING, 2 if any CRITICAL.
Query alert history from the SQLite database.
nerfwatch alerts # last 10 alerts
nerfwatch alerts --last 20 # last 20
nerfwatch alerts --severity CRITICAL # only critical
nerfwatch alerts --model claude-fable-5 # filter by model
nerfwatch alerts --json # structured outputManage the encrypted credential store. Nine subcommands cover the full lifecycle - enrolling a provider, listing what's stored, rotating keys, exporting the passphrase for backup, and migrating a v1 plaintext YAML into the v2 store.
nerfwatch auth login anthropic # Enroll a provider (prompts for key)
nerfwatch auth list # Which providers are currently enrolled
nerfwatch auth verify # Ping each provider to prove the stored key still works
nerfwatch auth rotate anthropic # Replace a key without restarting the monitor
nerfwatch auth remove anthropic # Tombstone a credential
nerfwatch auth revoke anthropic # Zero a credential and signal the running monitor
nerfwatch auth reset --confirm # Nuke the credstore (requires explicit confirmation)
nerfwatch auth migrate # One-shot import from v1 plaintext nerfwatch.yaml
nerfwatch auth export-passphrase # Print the passphrase for backup (mode 0600 recommended)Every subcommand routes through the atomic KeyProvider - a write flushes the credstore to disk, syncs the parent directory, and only then returns success. See docs/auth.md for threat model, rotation semantics, and the detailed failure-mode matrix.
Pin a single model to a higher grading tier for a bounded window of cycles. Typical use: a suspected regression on a Silver-tier model warrants Gold-tier validation for five cycles before deciding on further action. Think of it as ordering a stat echocardiogram when the routine monitor flags an anomaly - more evidence, same patient, no paging the full cardiac team.
nerfwatch burst --model claude-fable-5 --cycles 5 --tier gold
nerfwatch burst --cancel # Abort an active burst, no monitor restartburst writes ~/.nerfwatch/burst.json (mode 0600) signed with a per-install HMAC token at ~/.nerfwatch/.burst_token. The monitor verifies the signature every cycle and ignores a tampered file with a warn-level log. The web dashboard shows a BURST badge on the affected model card and a banner with the remaining cycles - cancel is idempotent and safe to run if you're not sure whether a burst is active.
Add a new model to an existing configuration.
nerfwatch configure add-modelInteractive prompt for provider, model ID, API key, and grader configuration. Validates that the grader model differs from the model under test.
Remove old monitoring data from the SQLite database.
nerfwatch purge --older-than 90d # dry run (default)
nerfwatch purge --older-than 90d --confirm # actually deleteProtects baseline data - refuses to purge records that overlap with active baseline computations.
Build the Docker image used by code probes. Required before code probes can execute.
nerfwatch sandbox buildThe image includes Python 3, Node.js (JavaScript/TypeScript via tsx), JDK (Java), g++ (C++), and Go. Requires Docker to be installed and running.
Check sandbox readiness - whether Docker is available and the sandbox image exists.
nerfwatch sandbox statusGenerate realistic demo data for the web dashboard. Creates 30 days of multi-model time-series data including regression events, gradual drift, and corresponding alerts.
nerfwatch seed # Generate demo data
nerfwatch seed --force # Regenerate (overwrites existing demo data)Uses synthetic model IDs with a demo- prefix (demo-claude-opus, demo-gpt-4o, demo-gemini-pro) that never collide with real monitoring data.
Manage Layer 3 frozen baselines for long-term drift detection.
nerfwatch anchor capture --model claude-fable-5 # Freeze current scores as baseline
nerfwatch anchor check --model claude-fable-5 # Re-run probes, compare to frozen baseline
nerfwatch anchor status # Display all anchor status and resultsView BOCPD-detected regime changes.
nerfwatch changepoints # Recent changepoints
nerfwatch changepoints --last 20 # Last 20
nerfwatch changepoints --model claude-fable-5 # Filter by model
nerfwatch changepoints --signal thinking_tokens # Filter by signalNerfWatch uses a hierarchical configuration system with three layers (highest priority first):
- Environment variables - nested overrides like
NERFWATCH_DB_PATH,NERFWATCH_SCHEDULE_INTERVAL, provider key overrides likeNERFWATCH_ANTHROPIC_API_KEY, and provider endpoint overrides likeNERFWATCH_ANTHROPIC_BASE_URL - YAML config file -
~/.nerfwatch/nerfwatch.yaml(or--configflag) - Compiled defaults - sensible defaults for all settings
Each model (and grader) accepts an optional base_url, and each provider
accepts a NERFWATCH_<PROVIDER>_BASE_URL env override (env wins). This routes
probes through any OpenAI-/Anthropic-compatible endpoint - for example a
local CLIProxyAPI sidecar that
re-exposes a flat-rate Claude or Codex subscription:
models:
- name: gpt-5.6-sol
provider: openai
model_id: gpt-5.6-sol
base_url: "http://localhost:8317" # host only - NerfWatch appends /v1/... itself
api_key_ref: "credentials://openai"Two rules, both enforced at config load:
base_urlmust be an absolutehttp(s)URL, host only - a trailing/v1or/v1betais rejected because the adapters append the version path themselves (a suffix would yield a doubled path and silent 404s).- Keep any local proxy bound to loopback; it is a gateway to a paid account.
Before trusting proxied samples in baselines, run the usage-parity smoke test
(TestSidecarUsageParity in internal/provider/) - if a proxy strips token
usage, its samples must never enter an SPC baseline. nerfwatch auth verify
always targets the provider's public endpoint regardless of base_url.
Credentials are never stored in the YAML file. They live in a passphrase-encrypted credstore at ~/.nerfwatch/credentials.enc (AES-256-GCM, Argon2id KDF) and the YAML references them by api_key_ref: pointers. This is the v2 format; legacy v1 configs with inline api_key: fields are migrated automatically on first nerfwatch init or explicitly via nerfwatch auth migrate. See docs/migration-v1-to-v2.md.
For reproducible installs, nerfwatch init --manifest consumes a small YAML manifest that lists the providers to enroll and the grader pairings to write. The manifest never contains secrets; keys are either prompted for interactively, piped in via --passphrase-stdin, or read from pre-populated environment variables (NERFWATCH_ANTHROPIC_API_KEY, etc.). See the manifest schema in docs/migration-v1-to-v2.md.
# v2 format: api_key_ref points into the encrypted credstore.
# Legacy v1 files with inline api_key: are auto-migrated on `nerfwatch init`.
models:
- name: claude-fable-5
provider: anthropic
model_id: claude-fable-5
api_key_ref: "credentials://anthropic"
schedule_interval: "30m" # Per-model override (expensive, run less often)
grader:
provider: openai
model_id: gpt-5.4-mini
api_key_ref: "credentials://openai"
- name: gpt-5.6-sol
provider: openai
model_id: gpt-5.6-sol
api_key_ref: "credentials://openai"
grader:
provider: anthropic
model_id: claude-haiku-4-5-20251001
api_key_ref: "credentials://anthropic"
- name: gemini-3.1-pro
provider: google
model_id: gemini-3.1-pro-preview
api_key_ref: "credentials://google"
grader:
provider: openai
model_id: gpt-5.4-mini
api_key_ref: "credentials://openai"
- name: grok-4.5
provider: xai
model_id: grok-4.5
api_key_ref: "credentials://xai"
grader:
provider: anthropic
model_id: claude-haiku-4-5-20251001
api_key_ref: "credentials://anthropic"
probe_file: probes.yaml
db_path: ~/.nerfwatch/nerfwatch.db
schedule:
interval: "15m" # Global default
alerts:
cooldown_hours: 4
jsonl_path: ~/.nerfwatch/alerts.jsonl
channels: # External alert delivery (optional)
- type: webhook
url: "https://example.com/nerfwatch-alerts"
- type: slack
url: "https://hooks.slack.com/services/T.../B.../..."
- type: email
smtp_host: smtp.gmail.com
smtp_port: 587
smtp_from: nerfwatch@example.com
smtp_to: ops@example.com
smtp_user: nerfwatch@example.com
smtp_pass: "app-password"
anchor:
check_interval: "168h" # Weekly re-check
alert_threshold: 0.10 # 10% delta triggers alertNerfWatch enforces that a model never grades its own output (a model that has degraded would rate its own degraded output as acceptable). Cross-provider grading is recommended: use an OpenAI model to grade Anthropic output and vice versa. Same-provider different-model grading (e.g., Haiku grading Opus) is permitted as a fallback.
NerfWatch labels every grading cycle with a tier based on enrolled providers (N) and peers-per-probe:
| Tier | N | peers/probe | Monitor behavior |
|---|---|---|---|
| Gold | ≥ 3 | N-1 | Every peer grades every target - strongest confidence. |
| Silver | ≥ 4 | 2 | Fixed two-peer grading - reduced cost, still cross-provider. |
| Bronze | ≥ 3 | 1 | One peer per probe - minimum viable cross-grading. |
| Deterministic-only | < 2 | - | LLM-judge disabled; falls back to exact / regex / numeric graders. Dashboard shows a red banner. |
When N drops below 3, the dashboard shows a cross-grading degraded banner so you know probe-quality confidence has fallen. Below N=2 the banner goes red and LLM-judge probes shut off entirely. Restore confidence by running nerfwatch auth login <provider> to enroll more. See docs/grading-tiers.md for the full tier math and promotion rules.
The example probes (probes.example.yaml) ship with known answers that models may have seen during training. For meaningful monitoring, create your own probes.yaml with domain-specific reasoning challenges.
Probe format:
version: "1.0.0"
probes:
- id: "my-logic-probe-01"
version: "1.0.0"
dimension: "logic" # logic, math, factual, or code
prompt: "Given {{.premise}}, what follows about {{.subject}}?"
variables:
premise: ["All X are Y", "No X are Y", "Some X are Y"]
subject: ["item A", "item B", "item C"]
answer_func: "logic-syllogism" # registered Go function
grading_method: "llm_judge" # exact, numeric_tolerance, regex, or llm_judgeEach run selects random values from the variable pools using a stored seed, so probes are reproducible but vary across runs to prevent memorization.
nerfwatch/
main.go # Entry point
cmd/ # Cobra CLI commands
root.go # Root command + global flags
init.go # nerfwatch init
run.go # nerfwatch run
monitor.go # nerfwatch monitor (main loop)
dashboard.go # nerfwatch dashboard (web + TUI)
status.go # nerfwatch status (one-liner)
alerts.go # nerfwatch alerts (history)
configure.go # nerfwatch configure
configure_addmodel.go # nerfwatch configure add-model
purge.go # nerfwatch purge
sandbox.go # nerfwatch sandbox (build/status)
seed.go # nerfwatch seed (demo data generation)
anchor.go # nerfwatch anchor (capture/check/status)
changepoints.go # nerfwatch changepoints (BOCPD results)
internal/
config/ # Hierarchical config (koanf)
storage/ # SQLite with WAL mode, dual pools
provider/ # Provider interface + adapters
provider.go # Request/Response types, Provider interface
anthropic.go # Anthropic adapter (ThinkingBlock extraction)
openai.go # OpenAI adapter (ReasoningTokens)
google.go # Google Gemini adapter (thinking parts)
xai.go # xAI Grok adapter (reasoning_content)
registry.go # Provider name -> implementation map
probe/ # Probe system
template.go # ProbeTemplate, GeneratedProbe, Generate()
loader.go # YAML loader with strict validation
answer_funcs.go # Answer function registry (8 MVP functions)
answer_funcs_code.go # Code probe answer functions (5 probes)
engine.go # Execution pipeline (sequential + concurrent grading)
grading/ # Grading system
deterministic.go # Exact, numeric tolerance, regex graders
llm_judge.go # LLM-as-judge with rubric + XML delimiters
sandbox.go # Docker sandbox grader for code probes
sandbox/ # Docker sandbox for code execution
sandbox.go # Container lifecycle (create, exec, teardown)
testspec.go # TestSpec type and JSON serialization
extractor.go # Code extraction from LLM responses
harness.go # Embedded per-language test runners
observe/ # Behavioral signal extraction
intrinsics.go # 5 Layer 1 signals from API responses
mews.go # MEWS composite score computation
stats/ # Statistical engines
ewma.go # EWMA tracker with time-varying control limits
cusum.go # CUSUM two-sided drift detection
bocpd.go # Bayesian Online Changepoint Detection
baseline.go # Per-model per-signal baseline management
anchor/ # Layer 3 frozen baselines
anchor.go # Capture, check, and status operations
alert/ # Alert system
alert.go # Evaluator with per-dimension cooldown
writer.go # JSONL file writer (mode 0600)
channels.go # Webhook, Slack, email delivery
server/ # Web dashboard backend
server.go # HTTP server with embedded SPA
api.go # REST API handlers (9 endpoints)
dashboard/ # Terminal UI
model.go # bubbletea Model (Init/Update/View)
views.go # lipgloss rendering + status formatting
web/ # React SPA (compiled, embedded via embed.FS)
probes.example.yaml # 13 example probes (3 logic, 3 math, 2 factual, 5 code)
nerfwatch.example.yaml # Example configuration
nerfwatch monitor
|
v
Probe Engine ──> Provider Adapter ──> AI Model API
| |
| Response + Usage
| |
+──> Behavioral Intrinsics Observer ──> Layer 1 signals
| |
| +──> MEWS Composite Score
|
+──> Grading System ──> Correctness + CoT quality
| |
| +──> Docker Sandbox (code probes)
|
+──> SQLite Storage (probe_runs, probe_results, intrinsic_observations)
|
v
Baseline Manager ──> EWMA Update ──> Threshold Check
| CUSUM Update ──> Drift Detection
| BOCPD ────────> Changepoint Detection
v
Alert Evaluator ──> Per-dimension cooldown check
|
+──> Internal: alerts.jsonl + SQLite alerts table
+──> External: Webhook + Slack + Email (configured channels)
|
v
Web Dashboard <── REST API (9 endpoints) <── SQLite read pool
NerfWatch uses SQLite with WAL mode for concurrent read/write access. The database lives at ~/.nerfwatch/nerfwatch.db by default.
- Write pool: single connection (
MaxOpenConns=1) for serialized writes - Read pool: 10 connections for dashboard, status, and alert queries
- Transactions: all writes use
BEGIN IMMEDIATEfor fast failure on contention - Schema: 9 tables (probe_runs, probe_results, intrinsic_observations, baselines, cusum_trackers, alerts, changepoints, anchors, anchor_results) with indices
NerfWatch uses three complementary detection engines from Statistical Process Control:
EWMA (Exponentially Weighted Moving Average) - catches sudden shifts:
- Lambda: 0.25 for Layer 2 probe dimensions, 0.2 for behavioral intrinsics
- Control limits: exact time-varying limits (wider during warm-up, converge to steady-state)
- Thresholds: WARNING at 2-sigma breach, CRITICAL at 3-sigma breach
- Calibration: 10 runs within a 7-day window required before alerting activates
CUSUM (Cumulative Sum) - catches sustained drift:
- K: 0.5 sigma (slack parameter)
- H: 4.0 sigma (decision boundary)
- Two-sided: independent upper and lower accumulators
- Fires DRIFT alerts when accumulated deviation exceeds H, independent of EWMA
BOCPD (Bayesian Online Changepoint Detection) - pinpoints regime changes:
- Identifies the exact moment a signal's behavior changes
- Records changepoints with confidence scores and affected dimensions
- Queryable via
nerfwatch changepoints
All three engines are direction-aware: some signals alert on decrease (thinking tokens), others on increase (inflation ratio), others bidirectionally (latency). A 73% drop in a signal (like the Opus regression) triggers EWMA CRITICAL within 1-3 observations and CUSUM DRIFT shortly after.
| Signal | What it measures | Direction | Provider support |
|---|---|---|---|
| ThinkingTokens | Chain-of-thought depth (word count proxy) | Below only | Anthropic (full), OpenAI (count only), Google (full), xAI (full) |
| InflationRatio | Output tokens / input tokens | Above only | All |
| LatencyProfile | Milliseconds per output token | Bidirectional | All |
| EarlyStopping | Self-interruption phrases ("Can I continue?") | Above only | All |
| ConnectorDensity | Logical connectors per 100 words in thinking | Below only | Anthropic, Google, xAI (not OpenAI) |
OpenAI models contribute 4 signals (no ConnectorDensity since reasoning content isn't exposed). Google and xAI provide full thinking/reasoning text, enabling all 5 signals like Anthropic. MEWS normalizes accordingly.
| NerfWatch | LangSmith | Patronus AI | |
|---|---|---|---|
| Approach | Standalone probes - zero code changes | SDK integration + tracing | SDK integration + guardrails |
| Detection | Statistical Process Control (EWMA + CUSUM + BOCPD) | Threshold-based evals | Rule-based + model evals |
| Catches | Unknown degradation you didn't anticipate | Known failures you define | Policy violations you specify |
| Metaphor | ICU cardiac monitor (always watching) | Lab blood test (you schedule it) | Security guard (checks rules) |
| Setup | nerfwatch init && nerfwatch monitor |
Instrument every LLM call | Instrument every LLM call |
| Cost | API keys only | Platform subscription | Platform subscription |
- Config file (
nerfwatch.yaml) and database created with mode0600 - Data directory (
~/.nerfwatch/) created with mode0700 - Alerts file (
alerts.jsonl) created with mode0600, contains derived metrics only - never raw model output - API keys are encrypted at rest in the credstore (
~/.nerfwatch/credentials.enc, AES-256-GCM + Argon2id); the YAML config stores only opaqueapi_key_ref:pointers. Env var override available (NERFWATCH_*_API_KEY). - Burst-window state file (
~/.nerfwatch/burst.json) is HMAC-signed with a per-install token so a tampered file is ignored by the monitor (warn-level log, no crash). - LLM judge input wrapped in XML delimiters to mitigate prompt injection from model-under-test output
Current main is feature-complete for a first public release: init + encrypted credentials, cross-grading engine, monitor with heartbeat and severity hysteresis, calibration report, web dashboard, seed command, full CI matrix (Go + Vitest + Playwright across chromium/firefox/webkit/mobile/visual/lighthouse). Telemetry deliberately declined for now - see docs/brainstorms/nerfwatch-telemetry-requirements.md.
What's left before tagging v0.1 is launch prep, not product work: README polish for a cold-reading first adopter, docs review pass, decision on tendaigomo/NerfWatch visibility flip, release tag + notes, and an adopter outreach plan.
- SDK wrapper integration mode (monitor production traffic without separate probes).
- Public hosted dashboard at nerfwatch.dev.
- User authentication and multi-tenant data isolation.
- WebSocket push (replace polling).
- Dashboard-driven configuration editing.
- Crash telemetry (Sentry-style error reporting - separate decision from the usage-telemetry rejection, needs its own brainstorm).
We welcome contributions! See CONTRIBUTING.md for guidelines on adding providers, probes, and statistical engines.
MIT - see LICENSE for details.




