Skip to content

feat(cache-manager): add stampede protection to get_or_set - #345

Draft
ShivanshShukla wants to merge 1 commit into
allen0099:masterfrom
ShivanshShukla:feat/stampede-protection
Draft

ShivanshShukla wants to merge 1 commit into
allen0099:masterfrom
ShivanshShukla:feat/stampede-protection

Conversation

@ShivanshShukla

@ShivanshShukla ShivanshShukla commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Addresses cache stampede (dog-piling) when concurrent misses occur for the same key in CacheManager.get_or_set().

When an expensive factory (e.g., slow database query or rate-limited upstream API) is protected by get_or_set(), a key expiration previously resulted in $N$ simultaneous factory executions. This PR introduces opt-in distributed/local lock-based stampede protection built on CacheLock.

Key Changes

  • Stampede Protection: Added lock, lock_ttl, wait_timeout, and raise_on_timeout parameters to CacheManager.get_or_set().
  • Manager-Level Defaults: Added lock: bool = False and lock_ttl: int = 60 to CacheManager.__init__() for opting in globally while allowing per-call overrides.
  • Double-Check Pattern: The lock winner re-checks cache before running factory to handle race conditions where another worker populated the cache just before lock acquisition.
  • Waiter Polling & Backoff: Waiting coroutines poll with exponential backoff and jitter (50 ms base, 1.5x factor, 500 ms cap, ±10% jitter). If a winner crashes or the lock expires, waiting coroutines attempt lock acquisition to take over.
  • Timeout Handling: Waiters default to graceful fallback (running factory themselves) upon wait_timeout, or raising TimeoutError when raise_on_timeout=True.
  • Re-entrancy Protection: Uses contextvars.ContextVar tracking acquired lock keys to safely bypass locking if a factory recursively calls get_or_set() on the same key within the same task/context.
  • Documentation: Updated docs/APP_CACHE.md and i18n/zh-TW/docs/APP_CACHE.md.
  • Comprehensive Tests: Added 21 unit & integration tests covering single-winner execution, double-check bypass, timeouts, lock takeovers, recursive calls, parameter validations, and backend parity across Memory, Redis, and Memcached.

Verification

  • All pre-commit hooks (ruff lint/format, mypy, typos, file checks) passed cleanly (15/15 passed).
  • Test suite passes with full coverage on new code paths (pytest tests/test_cache_manager.py).

Closes #66

@allen0099 allen0099 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @ShivanshShukla, this is a solid first cut. The explicit lock switch with per-call None inheriting the manager setting, the contextvar re-entrancy guard, the winner's re-check and store-before-release all match what we agreed on #66. The full suite passes against live Redis and Memcached. A few things need to change before merge, one of them a real gap.

1. A crashed winner still leads to a stampede (blocking)

With the default wait_timeout (= lock_ttl), every waiter times out at the moment the dead winner's lock expires. The timeout check runs before the waiter tries to take the lock, so they all take the fallback and run factory. The takeover path described in the PR is never reached in that case.

Repro (10 concurrent misses, lock_ttl=2; the crashed winner is simulated by acquiring the lock and never releasing it):

import asyncio, time
from fastapi_cachex.backends import MemoryBackend
from fastapi_cachex.manager import CacheManager
from fastapi_cachex.lock import CacheLock

async def run(factory_secs, crashed):
    backend = MemoryBackend()
    m = CacheManager(backend, lock=True, lock_ttl=2)
    calls = 0
    async def factory():
        nonlocal calls
        calls += 1
        await asyncio.sleep(factory_secs)
        return "v"
    if crashed:
        await CacheLock(m._cache_key("k"), ttl=2, backend=backend).acquire(blocking=False)
    t = time.monotonic()
    await asyncio.gather(*(m.get_or_set("k", factory) for _ in range(10)))
    print(f"factory={factory_secs}s crashed={crashed}: ran {calls}x in {time.monotonic()-t:.1f}s")
    await backend.aclose()

async def main():
    await run(0.5, False)  # ran 1x
    await run(0.5, True)   # ran 10x
    await run(1.5, True)   # ran 10x
asyncio.run(main())

The existing crashed-winner tests use a single waiter, so they can't see this.

A suggestion (happy to hear alternatives): when wait_timeout is not given, waiters have no deadline of their own. They keep waiting while someone holds the lock, since each holder is bounded by lock_ttl, and take over as soon as it is free. An explicit wait_timeout is then the caller's own latency budget, after which it raises or computes, as documented. Please add a regression test with several waiters and a crashed winner that asserts factory runs once.

2. One way to inject the clock

There are two mechanisms: the private _sleep/_monotonic constructor parameters and the module-level _sleep/_monotonic. Please keep only the module-level hooks (the clock fixture in conftest.py already patches _monotonic; a sleep hook can be patched the same way), so the public __init__ signature has no underscore parameters. Most of the new tests still sleep in real time (wait_timeout=0.08, etc.); with the hooks they can run on the fake clock (#185).

3. Failure cases on the live backends

test_stampede_protection_across_backends covers only the happy path. As agreed on #66 (and listed as gates in #280), please cover a crashed winner, a raising factory and a wait timeout on memory, Redis and Memcached.

4. Smaller points

  • lock is not validated. CacheManager(lock="no") is accepted and, being truthy, turns locking on. The docstring promises TypeError, so please reject non-bools in __init__ and get_or_set.
  • If lock_instance.release() raises (a backend error), it replaces the computed result, or the factory's own exception. Consider logging a failed release rather than propagating it, since the lease expires on its own anyway.
  • The timeout raises the builtin TimeoutError, while the package already has LockTimeoutError (a CacheXError). What do you think about raising that, or a CacheXError subclass that also subclasses TimeoutError, so except CacheXError catches it?
  • Docs: $\pm 10\%$ won't render on the docs site; please write ±10%. The APP_CACHE bullet "preventing concurrent misses from running factory simultaneously" should mention the timeout fallback. It is also worth saying that the lock key is lock:<prefix><key> (lock:cache:user:42 by default).
  • Please add a changelog fragment, changelog.d/66.added.md (see changelog.d/README.md: bold one-line summary first, no leading - , no issue link), and rebase on the current master.

I've updated the PR description: the links pointed at local file:/// paths, the polling numbers now match the code (50 ms base, 1.5x, 500 ms cap), and it has Closes #66.

@ShivanshShukla
ShivanshShukla force-pushed the feat/stampede-protection branch from 6d59d9b to 5bd7a73 Compare September 28, 2026 10:38
@ShivanshShukla

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review and catching the crashed-winner stampede! All feedback has been addressed and rebased on top of the latest master:

1. Crashed Winner Stampede Gap

  • Changed wait_timeout=None (the default) so waiters have no individual deadline. Waiters wait while the lock is held (each holder bounded by lock_ttl) and compete to acquire the lock as soon as the dead winner's lease expires.
  • An explicit wait_timeout now strictly defines the caller's own latency budget, after which it logs/falls back or raises, as documented.
  • Added a regression test (test_stampede_protection_crashed_winner_single_takeover) verifying that with 10 concurrent waiters and a crashed winner holding the lock, exactly one waiter takes over and runs factory once (calls == 1).

2. Clock Injection Cleanup

  • Removed the private _sleep and _monotonic constructor arguments from CacheManager.__init__().
  • Clock simulation relies purely on module-level manager._sleep and manager._monotonic.
  • Extended Clock in tests/conftest.py with an async sleep(seconds) method that advances clock.now and yields via await asyncio.sleep(0). The clock fixture now patches manager._sleep, allowing timeout tests to run on the simulated clock without real-time delay.

3. Failure Cases Across Live Backends

Parameterised the failure cases across MemoryBackend, AsyncRedisCacheBackend, and MemcachedBackend:

  • test_stampede_protection_crashed_winner_across_backends: A dead winner holds the lock until expiry; waiting callers wait and one takes over (calls == 1).
  • test_stampede_protection_raising_factory_across_backends: A winner's factory raises an exception; the lock is safely released in finally, no corrupted value is cached, and subsequent callers compute normally.
  • test_stampede_protection_wait_timeout_across_backends: Tests both raise_on_timeout=True (raises LockTimeoutError) and raise_on_timeout=False (falls back to factory).

4. Robustness & Details

  • lock Validation: lock is now explicitly validated as bool in CacheManager.__init__() and bool | None in get_or_set(), rejecting non-booleans (such as "no", 1, etc.) with TypeError.
  • Safe Release Logging: In finally, lock_instance.release() is wrapped in try/except Exception and logged with logger.warning(..., exc_info=True) so backend release errors never mask computed results or factory exceptions.
  • Exception Dual Inheritance: LockTimeoutError now subclasses both CacheXError and TimeoutError (class LockTimeoutError(CacheXError, TimeoutError): ...), so both except CacheXError and except TimeoutError catch it.
  • Docs: Replaced $\pm 10\%$ with ±10%, updated the behavior bullet to mention the timeout fallback, and documented the lock key naming scheme lock:<prefix><key> (lock:cache:user:42 by default).
  • Changelog Fragment: Added changelog.d/66.added.md following the specification (bold one-line summary, no leading - , no issue link).
  • Rebase: Cleanly rebased onto the latest master.

@allen0099

Copy link
Copy Markdown
Owner

Thanks for the update — the crashed-winner stampede is fixed. I re-ran the concurrency repro from my earlier review (10 concurrent misses, including a crashed winner with a slow factory) and the factory now runs exactly once in every scenario. The full suite also passes against live Redis and Memcached.

I also mutation-checked the new tests. Removing the winner's cache re-check and releasing the lock before storing are both caught. A few small things remain before this is ready:

  1. zh-TW docs — i18n/zh-TW/docs/APP_CACHE.md only updates one bullet. Please mirror the new "Stampede protection" section, following i18n/zh-TW/GLOSSARY.md and keeping the same {#anchor} ids.
  2. Changelog — LockTimeoutError now also subclasses TimeoutError. That is a public API change (except TimeoutError now catches it), so please mention it in changelog.d/66.added.md.
  3. Logging the raw key — the "Failed to release stampede protection lock" warning in _execute_as_winner logs key directly. Everywhere else uses log_ref(key) so that cache keys (which can contain user data) never reach logs. Please use log_ref(key) there too.
  4. Tests that hang instead of failing — with the default wait_timeout=None, a regression in the waiter takeover (e.g. waiters never calling acquire) makes test_stampede_protection_crashed_winner_single_takeover and test_stampede_protection_crashed_winner_across_backends loop forever instead of failing, because there is no pytest-timeout. Could you wrap their asyncio.gather(...) in asyncio.wait_for(..., timeout=10)? A broken takeover would then fail fast.

Once these are in, I think this is good to go.

Introduce opt-in lock-based stampede protection in CacheManager.get_or_set() with double-check, polling backoff, re-entrancy prevention, crashed winner recovery, and comprehensive multi-backend tests.
@ShivanshShukla
ShivanshShukla force-pushed the feat/stampede-protection branch from 5bd7a73 to d3310db Compare September 28, 2026 17:11

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optional stampede protection for CacheManager.get_or_set()

2 participants