A zero-dependency local HTTP proxy daemon that sits between an LLM-calling agent and its API endpoint. It records every request/response pair to disk as a JSON fixture, keyed by a normalized hash of the request's method, path, and body, so that test suites can later replay real traffic deterministically — offline, with no API keys and no network.
The proxy has two modes. In record mode it forwards requests to a
configurable upstream URL and passes the response back to the caller
unchanged. Chunked/SSE responses — the way LLM APIs stream tokens — are
forwarded live, chunk by chunk, as they arrive rather than buffered and
released all at once. Each chunk is also timestamped with its arrival delay
relative to the previous one. In replay mode no network connection is
made at all: each request is hashed the same way and looked up against the
fixtures directory. A match replays the recorded status, headers, and body,
re-emitting each chunk after its originally recorded delay so streaming pace
looks the same as the real run. A request with no matching fixture fails
loudly with a 500 and an x-llm-vcr-error: no-fixture header, rather than
silently passing through.
Requires Node.js 18+. No dependencies.
git clone <this-repo>
cd llm-vcr
npm installStart the proxy pointed at the real API you want to record. serve is the
default subcommand, so a bare --upstream flag still works too:
node bin/llm-vcr.js serve --upstream https://api.openai.com --port 8899 --fixtures-dir fixturesPoint your agent's OPENAI_BASE_URL (or equivalent) at
http://127.0.0.1:8899 instead of the real endpoint. Every request the agent
makes is forwarded to --upstream, the response is streamed back unmodified,
and the exchange is written to <fixtures-dir>/<hash>.json, where <hash> is
a sha256 digest of the request's method, path, and body.
For a streaming (SSE) response, the fixture's response also carries a
chunks array — one entry per chunk received from upstream, each with the
raw chunk text and delayMs, its arrival delay relative to the previous
chunk (or to the response starting, for the first one):
{
"response": {
"statusCode": 200,
"streaming": true,
"chunks": [
{ "data": "event: message\ndata: {\"delta\":\"Hel\"}\n\n", "delayMs": 42 },
{ "data": "event: message\ndata: {\"delta\":\"lo\"}\n\n", "delayMs": 18 }
],
"body": "..."
}
}Once fixtures exist, run the same test suite fully offline by switching to
replay mode — no --upstream, no API key, no network:
node bin/llm-vcr.js serve --replay --port 8899 --fixtures-dir fixturesEvery request is hashed the same way as at record time. A matching fixture
is replayed verbatim, chunk by chunk, waiting each chunk's originally
recorded delayMs before sending the next one. A request that doesn't match
any fixture (a typo, a code change, an argument that drifted) gets a clear
failure instead of a silent pass-through:
{
"error": "llm-vcr: no matching fixture for this request",
"method": "POST",
"path": "/v1/chat",
"key": "…sha256…"
}with x-llm-vcr-error: no-fixture on the response, so tests can assert on a
mismatch deliberately rather than mistaking it for a real API error.
serve options:
| Flag | Default | Description |
|---|---|---|
--upstream, -u |
(required in record mode) | Base URL of the real API to proxy to. |
--port, -p |
8899 |
Port the daemon listens on. |
--fixtures-dir, -f |
fixtures |
Directory fixtures are read from / written to. |
--mode, -m |
record |
record or replay. |
--replay |
Shorthand for --mode replay. |
Three more subcommands manage a fixtures directory once it has recordings in it — useful before committing fixtures to a repo, or cleaning out a stale cache in CI.
list — summarize every fixture (method, path, status, streaming shape, size):
node bin/llm-vcr.js list --fixtures-dir fixturesprune — delete fixtures recorded more than N days ago, based on each
fixture's own recordedAt timestamp. --dry-run reports what would be
deleted without touching anything:
node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 30 --dry-run
node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 30redact — strip secrets out of recorded fixtures before they're
committed: known auth header names (authorization, x-api-key,
openai-api-key, cookie, ...) are replaced outright, JSON body fields that
look like a credential (api_key, token, client_secret, ...) are
replaced outright, and any remaining text — non-JSON bodies, SSE chunk
data — is scanned for secret-shaped substrings (Bearer ..., sk-...,
AWS access key IDs) as a backstop for secrets under an unrecognized field
name. Only fixtures that actually contained something are rewritten:
node bin/llm-vcr.js redact --fixtures-dir fixtures --dry-run
node bin/llm-vcr.js redact --fixtures-dir fixturesAll three accept --fixtures-dir/-f (default fixtures), and only ever
touch files matching the <sha256-hex>.json shape that FixtureStore
writes — anything else in the directory is left alone.
By default a request is matched to a fixture by hashing its method, path, and body only — headers are never part of the hash, since auth headers and client-generated request IDs vary on every call without changing what the request actually asks for. That default needs no configuration and always applies.
Two situations still need to be handled explicitly, since they live inside what is hashed:
- A JSON request body sometimes carries a field that legitimately changes
every call — a
timestamp, anidempotency_key, a client-generatedmetadata.requestId— and without help, that field alone would make an otherwise-identical request fail to match its fixture on replay. - Rarely, a header genuinely does determine the response (an API version
header, say), and you want it to participate in matching — but the
volatile headers around it (
date,x-request-id,user-agent) still shouldn't.
Both are handled by a matching-rules JSON config file, passed with
--config/-c:
node bin/llm-vcr.js serve --replay --fixtures-dir fixtures --config llm-vcr.config.json{
"ignoreBodyFields": ["timestamp", "idempotency_key", "metadata.requestId"],
"matchHeaders": ["*"],
"ignoreHeaders": ["date", "x-request-id", "user-agent"]
}All three fields are optional and default to []:
| Field | Meaning |
|---|---|
ignoreBodyFields |
Dot-notation paths (e.g. "metadata.requestId") stripped out of a JSON request body before hashing. A missing path is a no-op, not an error. Non-JSON bodies are unaffected. |
matchHeaders |
Header names to fold into the hash. "*" means every header. Leave empty (the default) to keep headers out of matching entirely. |
ignoreHeaders |
Header names excluded even when matched by matchHeaders/"*" — this is how a header-matching wildcard stays useful despite the handful of headers that always vary. |
A config with an unrecognized field, a non-array field, or invalid JSON is rejected with a clear error rather than silently ignored — a typo here controls whether requests match too loosely or too strictly, so it fails loudly instead of producing confusing mismatches later.
The same rules apply on both record and replay, and to the
programmatic API:
import { startVcr } from 'llm-vcr';
const vcr = await startVcr({
fixturesDir: 'fixtures',
mode: 'replay',
configPath: 'llm-vcr.config.json', // or: matchRules: { ignoreBodyFields: [...] }
});Omitting --config/configPath entirely reproduces the exact matching
behavior llm-vcr has always had — existing fixtures keep matching without
any changes.
Programmatic use:
import { createProxyServer } from './src/index.js';
// record
const recorder = createProxyServer({
upstream: 'https://api.openai.com',
fixturesDir: 'fixtures',
});
recorder.listen(8899);
// replay — no upstream needed
const player = createProxyServer({ fixturesDir: 'fixtures', mode: 'replay' });
player.listen(8899);
// fixture management
import { listFixtures, pruneFixtures, redactFixturesDir } from './src/index.js';
const fixtures = await listFixtures('fixtures');
await pruneFixtures('fixtures', { maxAgeDays: 30 });
await redactFixturesDir('fixtures');src/testHarness.js starts and stops the daemon programmatically around a
Node test-runner suite, so node --test can run fully offline against
checked-in fixtures — no separate llm-vcr serve process, no network, no API
key.
The low-level pair mirrors server.listen()/server.close(), binding an
ephemeral port by default so parallel test files never collide:
import { startVcr, stopVcr } from 'llm-vcr';
let vcr;
before(async () => {
vcr = await startVcr({ fixturesDir: 'test/fixtures/agent', mode: 'replay' });
});
after(() => stopVcr(vcr));useVcr is sugar for the same thing: call it once at the top of a test file
to wire startVcr/stopVcr into that file's before/after hooks, and get
back a getter for the running daemon:
import { test } from 'node:test';
import assert from 'node:assert/strict';
import { useVcr } from 'llm-vcr';
const vcr = useVcr({ fixturesDir: 'test/fixtures/agent', mode: 'replay' });
test('agent gets a deterministic reply, offline', async () => {
const { baseUrl } = vcr();
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
method: 'POST',
body: JSON.stringify({ model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'hi' }] }),
});
assert.equal(res.status, 200);
});test/example.test.js is a full working copy of this pattern, complete with
a recorded fixture in test/fixtures/example-agent/ — read it as the
reference for wiring a real agent test suite up to llm-vcr.
The intended workflow is: record once, locally, with a real API key; commit the redacted fixtures; run the test suite in CI fully offline, with no key and no network egress at all.
-
Record locally, against the real upstream:
node bin/llm-vcr.js serve --upstream https://api.openai.com --fixtures-dir fixtures & OPENAI_BASE_URL=http://127.0.0.1:8899 npm test
-
Redact before committing — strip auth headers and credential-shaped fields out of every fixture, so no key ever reaches version control:
node bin/llm-vcr.js redact --fixtures-dir fixtures git add fixtures/ && git commit -m "record fixtures"
Run
redact --dry-runfirst if you want to see what would change without writing anything; see "Fixture management" above for exactly what counts as a secret. -
Replay in CI — no
--upstream, no key, no network. A typical GitHub Actions step:- run: node bin/llm-vcr.js serve --replay --fixtures-dir fixtures --port 8899 & - run: OPENAI_BASE_URL=http://127.0.0.1:8899 npm test
Or skip the standalone process entirely and use
useVcr/startVcr(see "Test-runner integration" above) so the daemon starts and stops with the test run itself — nothing to background, nothing to clean up. -
Keep fixtures matching real traffic. If a request legitimately includes a field that changes every run (a timestamp, an idempotency key), add a matching-rules config (see "Matching rules" above) rather than re-recording on every CI run — the point of committing fixtures is that CI never needs the real endpoint at all.
-
Prune stale fixtures periodically, so a repo doesn't accumulate fixtures for endpoints or prompts nobody tests against anymore:
node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 90
A request that doesn't match any committed fixture fails the build loudly
(500, x-llm-vcr-error: no-fixture) instead of silently reaching the real
API — the whole point of replay mode is that CI can't accidentally make a
live network call.
Built autonomously, gated on a passing test suite before any change is committed.
Run the tests:
npm test