Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llm-vcr

A zero-dependency local HTTP proxy daemon that sits between an LLM-calling agent and its API endpoint. It records every request/response pair to disk as a JSON fixture, keyed by a normalized hash of the request's method, path, and body, so that test suites can later replay real traffic deterministically — offline, with no API keys and no network.

The proxy has two modes. In record mode it forwards requests to a configurable upstream URL and passes the response back to the caller unchanged. Chunked/SSE responses — the way LLM APIs stream tokens — are forwarded live, chunk by chunk, as they arrive rather than buffered and released all at once. Each chunk is also timestamped with its arrival delay relative to the previous one. In replay mode no network connection is made at all: each request is hashed the same way and looked up against the fixtures directory. A match replays the recorded status, headers, and body, re-emitting each chunk after its originally recorded delay so streaming pace looks the same as the real run. A request with no matching fixture fails loudly with a 500 and an x-llm-vcr-error: no-fixture header, rather than silently passing through.

Install

Requires Node.js 18+. No dependencies.

git clone <this-repo>
cd llm-vcr
npm install

Usage

Start the proxy pointed at the real API you want to record. serve is the default subcommand, so a bare --upstream flag still works too:

node bin/llm-vcr.js serve --upstream https://api.openai.com --port 8899 --fixtures-dir fixtures

Point your agent's OPENAI_BASE_URL (or equivalent) at http://127.0.0.1:8899 instead of the real endpoint. Every request the agent makes is forwarded to --upstream, the response is streamed back unmodified, and the exchange is written to <fixtures-dir>/<hash>.json, where <hash> is a sha256 digest of the request's method, path, and body.

For a streaming (SSE) response, the fixture's response also carries a chunks array — one entry per chunk received from upstream, each with the raw chunk text and delayMs, its arrival delay relative to the previous chunk (or to the response starting, for the first one):

{
  "response": {
    "statusCode": 200,
    "streaming": true,
    "chunks": [
      { "data": "event: message\ndata: {\"delta\":\"Hel\"}\n\n", "delayMs": 42 },
      { "data": "event: message\ndata: {\"delta\":\"lo\"}\n\n", "delayMs": 18 }
    ],
    "body": "..."
  }
}

Once fixtures exist, run the same test suite fully offline by switching to replay mode — no --upstream, no API key, no network:

node bin/llm-vcr.js serve --replay --port 8899 --fixtures-dir fixtures

Every request is hashed the same way as at record time. A matching fixture is replayed verbatim, chunk by chunk, waiting each chunk's originally recorded delayMs before sending the next one. A request that doesn't match any fixture (a typo, a code change, an argument that drifted) gets a clear failure instead of a silent pass-through:

{
  "error": "llm-vcr: no matching fixture for this request",
  "method": "POST",
  "path": "/v1/chat",
  "key": "…sha256…"
}

with x-llm-vcr-error: no-fixture on the response, so tests can assert on a mismatch deliberately rather than mistaking it for a real API error.

serve options:

Flag Default Description
--upstream, -u (required in record mode) Base URL of the real API to proxy to.
--port, -p 8899 Port the daemon listens on.
--fixtures-dir, -f fixtures Directory fixtures are read from / written to.
--mode, -m record record or replay.
--replay Shorthand for --mode replay.

Fixture management

Three more subcommands manage a fixtures directory once it has recordings in it — useful before committing fixtures to a repo, or cleaning out a stale cache in CI.

list — summarize every fixture (method, path, status, streaming shape, size):

node bin/llm-vcr.js list --fixtures-dir fixtures

prune — delete fixtures recorded more than N days ago, based on each fixture's own recordedAt timestamp. --dry-run reports what would be deleted without touching anything:

node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 30 --dry-run
node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 30

redact — strip secrets out of recorded fixtures before they're committed: known auth header names (authorization, x-api-key, openai-api-key, cookie, ...) are replaced outright, JSON body fields that look like a credential (api_key, token, client_secret, ...) are replaced outright, and any remaining text — non-JSON bodies, SSE chunk data — is scanned for secret-shaped substrings (Bearer ..., sk-..., AWS access key IDs) as a backstop for secrets under an unrecognized field name. Only fixtures that actually contained something are rewritten:

node bin/llm-vcr.js redact --fixtures-dir fixtures --dry-run
node bin/llm-vcr.js redact --fixtures-dir fixtures

All three accept --fixtures-dir/-f (default fixtures), and only ever touch files matching the <sha256-hex>.json shape that FixtureStore writes — anything else in the directory is left alone.

Matching rules

By default a request is matched to a fixture by hashing its method, path, and body only — headers are never part of the hash, since auth headers and client-generated request IDs vary on every call without changing what the request actually asks for. That default needs no configuration and always applies.

Two situations still need to be handled explicitly, since they live inside what is hashed:

  • A JSON request body sometimes carries a field that legitimately changes every call — a timestamp, an idempotency_key, a client-generated metadata.requestId — and without help, that field alone would make an otherwise-identical request fail to match its fixture on replay.
  • Rarely, a header genuinely does determine the response (an API version header, say), and you want it to participate in matching — but the volatile headers around it (date, x-request-id, user-agent) still shouldn't.

Both are handled by a matching-rules JSON config file, passed with --config/-c:

node bin/llm-vcr.js serve --replay --fixtures-dir fixtures --config llm-vcr.config.json
{
  "ignoreBodyFields": ["timestamp", "idempotency_key", "metadata.requestId"],
  "matchHeaders": ["*"],
  "ignoreHeaders": ["date", "x-request-id", "user-agent"]
}

All three fields are optional and default to []:

Field Meaning
ignoreBodyFields Dot-notation paths (e.g. "metadata.requestId") stripped out of a JSON request body before hashing. A missing path is a no-op, not an error. Non-JSON bodies are unaffected.
matchHeaders Header names to fold into the hash. "*" means every header. Leave empty (the default) to keep headers out of matching entirely.
ignoreHeaders Header names excluded even when matched by matchHeaders/"*" — this is how a header-matching wildcard stays useful despite the handful of headers that always vary.

A config with an unrecognized field, a non-array field, or invalid JSON is rejected with a clear error rather than silently ignored — a typo here controls whether requests match too loosely or too strictly, so it fails loudly instead of producing confusing mismatches later.

The same rules apply on both record and replay, and to the programmatic API:

import { startVcr } from 'llm-vcr';

const vcr = await startVcr({
  fixturesDir: 'fixtures',
  mode: 'replay',
  configPath: 'llm-vcr.config.json', // or: matchRules: { ignoreBodyFields: [...] }
});

Omitting --config/configPath entirely reproduces the exact matching behavior llm-vcr has always had — existing fixtures keep matching without any changes.

Programmatic use:

import { createProxyServer } from './src/index.js';

// record
const recorder = createProxyServer({
  upstream: 'https://api.openai.com',
  fixturesDir: 'fixtures',
});
recorder.listen(8899);

// replay — no upstream needed
const player = createProxyServer({ fixturesDir: 'fixtures', mode: 'replay' });
player.listen(8899);

// fixture management
import { listFixtures, pruneFixtures, redactFixturesDir } from './src/index.js';

const fixtures = await listFixtures('fixtures');
await pruneFixtures('fixtures', { maxAgeDays: 30 });
await redactFixturesDir('fixtures');

Test-runner integration

src/testHarness.js starts and stops the daemon programmatically around a Node test-runner suite, so node --test can run fully offline against checked-in fixtures — no separate llm-vcr serve process, no network, no API key.

The low-level pair mirrors server.listen()/server.close(), binding an ephemeral port by default so parallel test files never collide:

import { startVcr, stopVcr } from 'llm-vcr';

let vcr;
before(async () => {
  vcr = await startVcr({ fixturesDir: 'test/fixtures/agent', mode: 'replay' });
});
after(() => stopVcr(vcr));

useVcr is sugar for the same thing: call it once at the top of a test file to wire startVcr/stopVcr into that file's before/after hooks, and get back a getter for the running daemon:

import { test } from 'node:test';
import assert from 'node:assert/strict';
import { useVcr } from 'llm-vcr';

const vcr = useVcr({ fixturesDir: 'test/fixtures/agent', mode: 'replay' });

test('agent gets a deterministic reply, offline', async () => {
  const { baseUrl } = vcr();
  const res = await fetch(`${baseUrl}/v1/chat/completions`, {
    method: 'POST',
    body: JSON.stringify({ model: 'gpt-4o-mini', messages: [{ role: 'user', content: 'hi' }] }),
  });
  assert.equal(res.status, 200);
});

test/example.test.js is a full working copy of this pattern, complete with a recorded fixture in test/fixtures/example-agent/ — read it as the reference for wiring a real agent test suite up to llm-vcr.

CI usage

The intended workflow is: record once, locally, with a real API key; commit the redacted fixtures; run the test suite in CI fully offline, with no key and no network egress at all.

  1. Record locally, against the real upstream:

    node bin/llm-vcr.js serve --upstream https://api.openai.com --fixtures-dir fixtures &
    OPENAI_BASE_URL=http://127.0.0.1:8899 npm test
  2. Redact before committing — strip auth headers and credential-shaped fields out of every fixture, so no key ever reaches version control:

    node bin/llm-vcr.js redact --fixtures-dir fixtures
    git add fixtures/ && git commit -m "record fixtures"

    Run redact --dry-run first if you want to see what would change without writing anything; see "Fixture management" above for exactly what counts as a secret.

  3. Replay in CI — no --upstream, no key, no network. A typical GitHub Actions step:

    - run: node bin/llm-vcr.js serve --replay --fixtures-dir fixtures --port 8899 &
    - run: OPENAI_BASE_URL=http://127.0.0.1:8899 npm test

    Or skip the standalone process entirely and use useVcr/startVcr (see "Test-runner integration" above) so the daemon starts and stops with the test run itself — nothing to background, nothing to clean up.

  4. Keep fixtures matching real traffic. If a request legitimately includes a field that changes every run (a timestamp, an idempotency key), add a matching-rules config (see "Matching rules" above) rather than re-recording on every CI run — the point of committing fixtures is that CI never needs the real endpoint at all.

  5. Prune stale fixtures periodically, so a repo doesn't accumulate fixtures for endpoints or prompts nobody tests against anymore:

    node bin/llm-vcr.js prune --fixtures-dir fixtures --older-than 90

A request that doesn't match any committed fixture fails the build loudly (500, x-llm-vcr-error: no-fixture) instead of silently reaching the real API — the whole point of replay mode is that CI can't accidentally make a live network call.

Status

Built autonomously, gated on a passing test suite before any change is committed.

Run the tests:

npm test

About

A zero-dependency local HTTP proxy daemon that sits between an LLM-calling agent and its API endpoint, recording every request/response (including SSE streaming chunks with original timing) to disk…

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages