Ziggurat models your cache as an ordered stack of layers. Each layer is a CacheAdapter — a storage backend implementing the full adapter interface.
┌─────────────────────┐
│ L1: MemoryAdapter │ ← Fastest (in-process, ~0ms)
├─────────────────────┤
│ L2: RedisAdapter │ ← Shared (network, ~1-5ms)
├─────────────────────┤
│ L3: Your Adapter │ ← Any CacheAdapter implementation
└─────────────────────┘
Layers are passed as an ordered array to the CacheManager. The first layer is the fastest and most local; deeper layers are progressively slower but may be shared across processes or machines.
const cache = new CacheManager({
layers: [l1, l2, l3], // checked in this order
});When you call cache.get(key) or cache.wrap(key, factory), the CacheManager queries layers sequentially, starting from L1:
- Check L1. If found → return the value.
- Check L2. If found → return the value and backfill L1.
- Check L3. If found → return the value and backfill L1 and L2.
- If no layer has the value → return
null(or call the factory forwrap).
This means the first layer to contain the value wins. Deeper layers are only queried on a miss.
The namespace option on CacheManager prefixes all keys with namespace:. This provides logical grouping without polluting your key strings:
const userCache = new CacheManager({
namespace: "users",
layers: [memory, redis],
});
// cache.set("42", data) → stored as "users:42"
// cache.get("42") → looks up "users:42"Different CacheManagers with different namespaces sharing the same adapters won't collide — users:42 and products:42 are distinct keys.
Namespaces are joined with : and not escaped: namespace "a" + key "b:c" produces the same stored key as namespace "a:b" + key "c". Avoid : in namespace values if your keys may contain it.
When a value is found in a lower layer (e.g., L2), the CacheManager automatically backfills all higher layers (e.g., L1) so subsequent reads are served from the fastest layer.
Each target layer applies its own TTL policy to the copy it receives, capped by whatever life the source entry has left:
- Target has a
defaultTtlMs→ the copy getsmin(defaultTtlMs, source's remaining TTL). - Target has no
defaultTtlMs→ the copy gets the source's remaining TTL. - Source entry is permanent → the target applies its own policy with nothing to cap it.
- A target's
maxTtlMscaps the result either way.
So L1 keeps its own staleness budget rather than inheriting L2's, and a backfilled copy never outlives the entry it was copied from. In a three-layer stack each target is computed independently — an L3 hit can backfill L1 with 10s and L2 with 60s from the same read.
By default, backfill is asynchronous — the value is returned immediately while backfill happens in the background. This minimizes response latency.
// Default: async backfill (fire-and-forget)
const cache = new CacheManager({
layers: [l1, l2],
});Set syncBackfill: true to wait for backfill to complete before returning. This guarantees that higher layers are populated before your code continues, which can be useful for testing or consistency-sensitive scenarios.
// Sync backfill: waits for L1 to be populated before returning
const cache = new CacheManager({
layers: [l1, l2],
syncBackfill: true,
});set and delete operate on all layers simultaneously using Promise.allSettled. Cross-layer operations are best-effort, not atomic — there is no distributed transaction across independent backends. If one layer fails, the operation still completes on the others, which may leave layers temporarily inconsistent until the next write reconciles them.
// Writes to L1 AND L2
await cache.set("key", value, 60_000);
// Deletes from L1 AND L2
await cache.delete("key");The CacheManager orchestrates keyed operations (get, set, delete, wrap, mget, mset, mdel, has, getTtl) across layers. Bulk backend operations like clear(), flushAll(), and keys() are adapter concerns — they operate on a single backend with well-defined scope.
Use getLayers() to access adapters directly when you need these operations:
const [l1, l2] = cache.getLayers();
// Clear a specific layer
await l1.clear();
// Clear all layers explicitly (you control the scope)
for (const layer of cache.getLayers()) {
await layer.clear();
}
// Get keys from a specific layer
const keys = await l2.keys();Adapters returned by getLayers() do not apply the manager's namespace — they expose their raw API. The array is a new copy on each call; mutations to it do not affect the manager.
When a popular cache key expires, many concurrent requests may all experience a cache miss at the same time. Without protection, all of them hit the underlying data source simultaneously — a cache stampede (also called "thundering herd" or "dogpile").
Request 1 → cache miss → DB query ─┐
Request 2 → cache miss → DB query ─┤
Request 3 → cache miss → DB query ─┤ All hit the DB at once!
... │
Request N → cache miss → DB query ─┘
Ziggurat uses request coalescing (in-flight deduplication). When the first request triggers a cache miss and calls the factory function, all subsequent requests for the same key attach to the existing in-flight Promise instead of creating new ones.
Request 1 → cache miss → factory() ──→ result ──→ cache + return
Request 2 → cache miss → [coalesced] ─────────→ same result
Request 3 → cache miss → [coalesced] ─────────→ same result
...
Request N → cache miss → [coalesced] ─────────→ same result
The factory executes exactly once. All N callers get the same value.
Coalescing is enabled by default. You can disable it for debugging or specific use cases:
const cache = new CacheManager({
layers: [new MemoryAdapter()],
stampede: { coalesce: false },
});If the factory function throws an error during coalescing, the error propagates to all coalesced callers. The in-flight entry is cleaned up so subsequent calls trigger a fresh factory invocation.
TTL is specified in milliseconds. There are two ways to configure it:
Set defaultTtlMs on each adapter — the fallback used when a call passes no TTL of its own. This is the recommended approach for multi-layer setups because each layer can have its own expiration policy:
const cache = new CacheManager({
layers: [
new MemoryAdapter({ defaultTtlMs: 30_000 }), // L1: 30s
new RedisAdapter({ client: redis, defaultTtlMs: 600_000 }), // L2: 10min
],
});
// No TTL needed in wrap — adapters handle it
await cache.wrap("key", factory);You can also pass a TTL directly to set or wrap. An explicit TTL wins over the adapter's defaultTtlMs:
// Overrides defaultTtlMs on every layer
await cache.set("key", value, 300_000);
await cache.wrap("key", factory, 300_000);- TTL passed via
set/wrap(if given) — wins - Adapter's
defaultTtlMs— fallback when no TTL was passed - No TTL — entry never expires
The adapter's maxTtlMs, when set, caps whatever the first two steps produce — including otherwise-permanent entries. Use defaultTtlMs for "what a layer does when nobody says otherwise" and maxTtlMs for "what a layer will never exceed."
Internally, TTL is stored as an absolute Unix timestamp (expiresAt) on the CacheEntry. Expired entries are lazily cleaned up on the next get call.
set(key, undefined) and mset entries whose value is undefined are silently skipped by every adapter — no backend can round-trip undefined, so storing it would read back as either a miss or a hit carrying undefined, depending on the layer. The write is a no-op: get, has, getTtl, and keys all report the key as absent, and an existing value under that key is left untouched rather than being overwritten.
The CacheManager emits typed events for every operation — hits, misses, errors, backfills, stampede coalescing, and more. Events have zero cost when no listeners are attached.
const cache = new CacheManager({
layers: [memory, redis],
});
// Log cache misses
cache.on("miss", (e) => {
console.log(`Miss: ${e.key} (${e.durationMs.toFixed(1)}ms)`);
});
// Track errors per layer
cache.on("error", (e) => {
console.error(`Layer ${e.layerName} failed on ${e.operation}:`, e.error);
});
// Monitor backfill activity
cache.on("backfill", (e) => {
console.log(
`Backfill: ${e.key} from ${e.sourceLayerName} → ${e.targetLayerNames.join(", ")}`,
);
});The on() method returns an unsubscribe function:
const unsub = cache.on("hit", listener);
// ...later
unsub(); // stop listeningThe @ziggurat-cache/otel package translates cache events into OTel counters and histograms. It only depends on @opentelemetry/api (the lightweight API, not the SDK) — your application provides the SDK and exporter.
import { instrumentCacheManager } from "@ziggurat-cache/otel";
const cleanup = instrumentCacheManager(cache);
// Metrics are now flowing to your OTel backendSee the API Reference for the full list of recorded metrics.
Ziggurat is designed to be resilient. Cross-layer operations are best-effort, not atomic — there is no distributed transaction across independent backends. Individual layer failures never crash the overall operation:
get: If a layer throws during a read, it is skipped and the next layer is queried.set/delete: Operations usePromise.allSettled, so failures on one layer don't prevent success on others. This can leave layers temporarily inconsistent until the next write reconciles them.- Backfill failures: If backfilling a higher layer fails after a lower-layer hit, the value is still returned to the caller.
This means a transient Redis connection error won't prevent your application from serving data from memory — or vice versa.