Skip to content

feat: add CacheLock and the expire_if_equals backend primitive - #150

Merged
allen0099 merged 9 commits into
allen0099:masterfrom
ShivanshShukla:feat/cache-lock
Sep 26, 2026
Merged

allen0099 merged 9 commits into
allen0099:masterfrom
ShivanshShukla:feat/cache-lock

Conversation

@ShivanshShukla

@ShivanshShukla ShivanshShukla commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #64. The expire_if_equals primitive was tracked in #62.

This PR introduces:

  1. An owner-checked TTL renewal primitive (expire_if_equals) across all backends.
  2. A distributed CacheLock helper usable as an async context manager or via explicit acquire/release/extend/locked calls.

Proposed Changes

1. expire_if_equals(key, expected, ttl) -> bool Primitive

  • BaseCacheBackend: Added non-abstract method with non-atomic fallback and validate_ttl(ttl).
  • MemoryBackend: Overridden under self.lock.
  • AsyncRedisCacheBackend: Overridden using Lua script _EXPIRE_IF_EQUALS_SCRIPT (GET compare + EXPIRE).
  • MemcachedBackend: Overridden using GETS + CAS write with the new exptime (TOUCH takes no CAS token).
  • docs/BACKENDS.md: Documented expire_if_equals under the Atomic backend primitives section.

2. CacheLock Distributed Lock & Helpers

  • CacheLock (fastapi_cachex/lock.py):
    • acquire(blocking=True, timeout=None, poll_interval=0.1, ttl=None) -> bool: non-blocking mode uses single set_if_absent; blocking mode retries with deadline loop until timeout.
    • release() -> bool: calls delete_if_equals. Silently returns False if lock was lost or expired.
    • extend(ttl=None) -> bool: calls expire_if_equals to safely renew holder's lease.
    • locked() -> bool: queries whether lock key exists in backend.
    • Configurable key_prefix="lock:" (defaults under lock: namespace).
    • Explicit backend= parameter defaulting to BackendProxy.get().
    • Unique holder token per instance generated via secrets.token_hex(16).
  • LockTimeoutError (fastapi_cachex/exceptions.py): Raised by __aenter__ when context manager async with CacheLock(..., timeout=X) fails to acquire lock within timeout.
  • Root Exports: Exported CacheLock and LockTimeoutError from fastapi_cachex root.
  • CHANGELOG.md: Added release notes under [Unreleased].

Usage Example

from fastapi import HTTPException
from fastapi_cachex import CacheLock, LockTimeoutError

# As an async context manager (raises LockTimeoutError on timeout)
async with CacheLock("report:123", ttl=30):
    ...  # exclusive execution across processes

# Direct call pattern
lock = CacheLock("stream:user_42", ttl=60)
if not await lock.acquire(blocking=False):
    raise HTTPException(409, detail="Lock already held")
try:
    ...
    await lock.extend(60)  # renew lock before expiration
finally:
    await lock.release()

…ext manager (allen0099#62)

- Implement expire_if_equals(key, expected, ttl) atomic primitive on BaseCacheBackend, MemoryBackend, AsyncRedisCacheBackend (Lua script), and MemcachedBackend (CAS).
- Document expire_if_equals in docs/BACKENDS.md under 'Atomic backend primitives'.
- Implement CacheLock distributed lock helper and async context manager in fastapi_cachex/lock.py.
- Add LockTimeoutError exception raised on context manager timeout.
- Export CacheLock and LockTimeoutError from package root.
- Add comprehensive unit test suite in tests/test_lock.py and tests/backends/.
- Add CHANGELOG entry under [Unreleased].
@ShivanshShukla

Copy link
Copy Markdown
Contributor Author

Hi @allen0099,
I've raised a draft for you to review. Looking forward to it.

@allen0099 allen0099 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, @ShivanshShukla, this is a solid first cut, and it follows the design we agreed on in #64 closely. acquire() returns False and only __aenter__ raises LockTimeoutError, a lost release() stays quiet, and the Redis and Memcached expire_if_equals implementations mirror delete_if_equals, including the race tests. A few things before this can leave draft:

Scope and linking

  1. Let's keep this as one PR rather than splitting it. Both halves are already here and belong together. Please add Closes #64 to the description so the issue closes on merge. #62 is the earlier, already closed primitive issue.

Code

  1. fastapi_cachex/lock.py imports Self from typing_extensions, which is not a declared dependency of this package (it only arrives through pydantic). Please return "CacheLock" from __aenter__ instead, as a quoted annotation, the way the rest of the package handles forward references.

  2. LOCK_FINGERPRINT and lock_entry() in types.py become public API, but only lock.py uses them. Please move them into lock.py as private helpers (_LOCK_FINGERPRINT, _lock_entry).

  3. acquire(timeout=None) means "use the instance default", so a caller can't ask for an unbounded wait on one call when the instance has a timeout. That's acceptable, but please say so in the acquire docstring.

  4. Blocking: the token belongs to the instance, so two acquisitions through the same instance share it, and that silently defeats the owner check. Two ways to get there:

    • Reentry. Calling acquire() again on an instance that already holds the lock polls until the lock's own TTL expires. It then re-acquires with the same token while the outer critical section is still running. When the inner block releases, delete_if_equals matches and frees the lock under the outer block, so another process can take it while the outer code still runs. Nothing is raised or logged.
    • A shared instance. This one is more likely, because it is how people use asyncio.Lock: a module-level report_lock = CacheLock("report") used with async with report_lock: in concurrent requests. Once one request overruns the TTL and another takes the lock, the first request's release() deletes the second request's lock. That is exactly the race delete_if_equals exists to prevent.

    Documentation alone won't stop either of these. Please track holding state on the instance and make acquire() raise RuntimeError when the instance already holds the lock. The message should tell the caller to use one CacheLock per acquisition. Clear the state on release() (successful or not). Please don't return False here: False means "someone else holds it", and under async with it would surface as a misleading LockTimeoutError. Also mention the one-instance-per-acquisition rule in the class docstring, and add tests for both the reentrant and the shared-instance case.

Tests

  1. CacheLock itself is only exercised against MemoryBackend. Please add at least one acquire → extend → release round trip against Redis and Memcached (requires_redis / requires_memcached), since that is where the lock matters.
  2. Nit: move the import asyncio in test_lock_acquire_blocking_indefinite_retries_until_available to the top of the module.
  3. CI is green on your side: all tests pass on 3.10–3.14 and in tox, and coverage stays at 100% with live Redis and Memcached. The red "coverage" check comes from its badge-upload step, which tries to push to this repo and can't from a fork. That's a bug in our workflow, not in your PR, and I'll fix it separately. Please keep coverage at 100% as you add the changes above.

Docs

  1. CacheLock has no user-facing docs yet. Please add:
    • a new guide, docs/LOCK.md, added to the nav in zensical.toml next to the other guides. It should cover the two usage patterns from #64, blocking vs. non-blocking acquire and timeouts, renewing with extend(), what happens when the TTL runs out before release(), and the one-instance-per-acquisition rule from point 5. Please also add it to the documentation list in README.md;
    • an API reference entry, e.g. ::: fastapi_cachex.lock.CacheLock in a new docs/api/lock.md added to the nav in zensical.toml. LockTimeoutError already appears in docs/api/types.md, which documents the whole exceptions module.
  2. Please link #64 in the CHANGELOG entries (([#64](https://github.com/allen0099/FastAPI-CacheX/issues/64))), as the other entries do.

You don't need to touch the Traditional Chinese docs under i18n/. Those get translated separately.

Thanks again!

allen0099 added a commit that referenced this pull request Sep 25, 2026
The Coverage Badge workflow runs on pull requests too, and its last step
pushed the badge to the coverage-badge branch every time. A pull request
from a branch in this repository overwrote master's badge with the pull
request's coverage. A pull request from a fork gets a read-only token, so
the push failed with 403 and the check went red even though tests and
coverage passed (seen on #150).

Run the upload step only for push events. Pull requests still run the
suite against live servers and enforce the coverage gate.
ShivanshShukla and others added 2 commits September 25, 2026 22:25
)

- Move _LOCK_FINGERPRINT and _lock_entry into lock.py as private helpers.
- Add _is_held state guard to CacheLock to prevent re-entry or instance sharing across concurrent tasks.
- Remove Self import from typing_extensions; use string return annotation for __aenter__.
- Update acquire() docstring timeout parameter description.
- Add RuntimeError re-entry and shared instance unit tests in tests/test_lock.py.
- Add live Redis and Memcached integration tests for CacheLock.
- Add docs/LOCK.md guide, docs/api/lock.md API reference, and update navigation in zensical.toml, zensical.zh-TW.toml, and README.md.
- Update CHANGELOG.md entry with issue link allen0099#64.
@ShivanshShukla

Copy link
Copy Markdown
Contributor Author

@allen0099 Thanks for the thorough review! The catch on the shared instance token overlap was spot on.

I've pushed a new commit addressing everything:
Added _is_held state tracking to prevent reentry/shared-instance usage, raising RuntimeError as requested, with tests for both scenarios.
Added Redis and Memcached round-trip tests for the lock lifecycle.
Moved the helpers to lock.py as private, fixed the Self typing, and cleaned up the asyncio import.
Wrote the docs/LOCK.md guide and API reference, and wired them into zensical.toml and the README.
Added Closes #64 to the PR description and CHANGELOG.
Let me know if the docs or the RuntimeError implementation need any further tweaking!

@allen0099 allen0099 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the quick turnaround, @ShivanshShukla. Most of the first review is addressed: the private helpers, the quoted __aenter__ annotation, the acquire docstring, the live Redis and Memcached round trips, and the new guide and API page all look good. A few things are still left:

Blocking

  1. The shared-instance check can still be bypassed. acquire() only sets _is_held after set_if_absent succeeds, and there is an await in between. Two tasks that call acquire() on the same instance while it is free both get past the check. One of them wins, and the other keeps polling with the same token. If the winner overruns the TTL, the loser takes the lock with that token, and the winner's release() deletes it: the original race. test_lock_shared_instance_raises_runtime_error passes only because MemoryBackend.set_if_absent doesn't yield when its lock is free. Over Redis or Memcached, where every call really awaits, both tasks pass the check.

    I reproduced it against a live Redis with one shared CacheLock("shared", ttl=1): task A holds it for 1.5 s, task B calls acquire() at the same time, and a separate CacheLock("shared") tries a non-blocking acquire while B is inside. Nothing raises, and the log is:

    A acquired
    B acquired                  # while A is still inside
    A released -> True          # deletes the lock B now holds
    outsider acquired -> True   # while B is still inside
    B released -> False
    

    Please mark the instance as in use at the top of acquire(), before the first await, and clear the mark if the acquisition fails or times out (and in release(), as now). A second acquire() on the instance then raises RuntimeError right away, whatever the backend does. For the test, please use a backend whose set_if_absent yields (for example, a MemoryBackend subclass that does await asyncio.sleep(0) first), or run the shared-instance case against Redis as well, so the test fails without the fix.

  2. CHANGELOG format. #153 has just been merged: every changelog entry now has to open with a bold one-line summary, because the GitHub release notes are built from those summaries. Your merge from master already brought it in, so CI will fail on this branch as it stands (test_the_repository_changelog_can_be_released). Please reword both entries along these lines, link #64 on the expire_if_equals entry too, and leave a blank line before ### Changed:

    ### Added
    
    - **`CacheLock`, a distributed lock built on the backend primitives.** ...details... ([#64](https://github.com/allen0099/FastAPI-CacheX/issues/64))
    - **`expire_if_equals()` backend primitive for owner-checked TTL renewal.** ...details... ([#64](https://github.com/allen0099/FastAPI-CacheX/issues/64))

    See "The changelog is part of the release now" in docs/DEVELOPMENT.md.

Docs

  1. Please revert the change to zensical.zh-TW.toml. The Traditional Chinese site has no LOCK.md, so the new nav entry links to a 404 in the preview (https://fastapi-cachex--150.org.readthedocs.build/zh-tw/150/LOCK/). The page gets added there when it is translated.
  2. docs/LOCK.md doesn't yet cover what happens when the TTL runs out before release(). Please say that plainly: the lock becomes free, another process can take it while your code is still running, and extend()/release() then return False. Advise choosing a TTL longer than the work, or calling extend() periodically. Please also say that the default timeout=None makes a blocking acquire() (and async with) wait indefinitely.

Minor

  1. Instead of a file-wide PYI034 ignore in pyproject.toml, please use # noqa: PYI034 on the __aenter__ line, so the rule stays active for the rest of the module.
  2. The PR description still says it addresses #62, and it doesn't have Closes #64 yet. Please update it so the issue closes on merge.

Thanks!

…len0099#64)

- Set _is_held = True synchronously at entry of acquire() before the first await to prevent shared-instance race conditions.
- Format CHANGELOG.md entries with bold summaries and link allen0099#64 on both.
- Revert zensical.zh-TW.toml edit.
- Update docs/LOCK.md to detail TTL expiration behavior and default timeout=None indefinite blocking.
- Move PYI034 ignore inline on __aenter__ line in lock.py.
- Update test_lock_shared_instance_raises_runtime_error in tests/test_lock.py to use YieldingMemoryBackend.
@allen0099 allen0099 changed the title feat: add expire_if_equals backend primitive and CacheLock async cont… feat: add CacheLock and the expire_if_equals backend primitive Sep 25, 2026

@allen0099 allen0099 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, @ShivanshShukla, this round looks good. I re-ran my shared-instance reproduction against a live Redis: the second acquire() now raises RuntimeError straight away, and the race is gone. The CHANGELOG entries, the LOCK.md additions, the inline noqa and the zensical.zh-TW.toml revert are all as asked. One thing left:

A cancelled acquire() leaves the instance stuck. acquire() clears _is_held in except Exception, but asyncio.CancelledError is a BaseException, not an Exception. So a blocking acquire() that is cancelled, for example by asyncio.wait_for()/asyncio.timeout() or a cancelled request, leaves _is_held = True, and every later acquire() on that instance raises RuntimeError even though it never got the lock:

waiter = CacheLock("job", ttl=30)          # "job" is held elsewhere
try:
    await asyncio.wait_for(waiter.acquire(), timeout=0.3)
except asyncio.TimeoutError:
    pass
waiter._is_held                            # True
await waiter.acquire(blocking=False)       # RuntimeError: ... already held

Please change it to except BaseException: (it re-raises, so nothing is swallowed), and add a test that cancels a blocking acquire() and then acquires again on the same instance.

I've updated the PR title (it was cut off) and the description so that it says Closes #64, which will close the issue on merge. No action needed there.

After that, please mark the PR ready for review.

@allen0099

Copy link
Copy Markdown
Owner

One more thing, which CI found once it ran: test_lock_lifecycle_with_redis and test_lock_lifecycle_with_memcached error with fixture 'async_redis_backend' not found / fixture 'memcached_backend' not found. Those fixtures are defined in tests/backends/test_redis.py and tests/backends/test_memcached.py, not in a conftest.py, so tests/test_lock.py can't see them. They skip locally without live servers, which is why this doesn't show up there, and I missed it the same way.

The simplest fix is to move the two tests into tests/backends/test_redis.py and tests/backends/test_memcached.py, next to the fixtures they use. Everything else passes (787 tests).

@ShivanshShukla

Copy link
Copy Markdown
Contributor Author

Thanks for catching that, @allen0099!
That makes total sense — since async_redis_backend and memcached_backend are defined locally within tests/backends/test_redis.py and tests/backends/test_memcached.py, moving the integration tests into their respective backend files fixes the fixture lookup errors on CI:

  • Moved test_lock_lifecycle_with_redis to tests/backends/test_redis.py next to async_redis_backend.
  • Moved test_lock_lifecycle_with_memcached to tests/backends/test_memcached.py next to memcached_backend.
  • Cleaned up tests/test_lock.py by removing unused live-server decorator imports and the unreferenced TYPE_CHECKING block.
    All 625 unit tests, ruff, mypy --strict, and changelog validation checks pass cleanly.

Lint failed on ruff format --check for the two CacheLock lifecycle tests moved into tests/backends/.
@allen0099

allen0099 commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Thanks, the fixture fix works: with live servers in CI, both lifecycle tests now pass.

Lint failed only because ruff format --check flagged a trailing blank line at the end of tests/backends/test_redis.py and tests/backends/test_memcached.py. I pushed 025b552 to your branch to remove them, so please git pull before your next push.

One point from my review just before the fixture comment is still open: acquire() still clears _is_held in except Exception:. An asyncio.CancelledError is a BaseException, so a cancelled acquire() leaves the instance stuck. That review has the reproduction. Please change it to except BaseException: and add a test that cancels a blocking acquire() and then acquires again on the same instance. After that, this is ready to leave draft.

@ShivanshShukla

Copy link
Copy Markdown
Contributor Author

Thanks @allen0099! Pulled 025b552 and addressed the cancellation point:

  • Updated CacheLock.acquire() Exception Handler: Changed except Exception: to except BaseException: in fastapi_cachex/lock.py so asyncio.CancelledError properly resets self._is_held = False when a task waiting in acquire() is cancelled.
  • Added Cancellation Unit Test: Added test_lock_acquire_cancelled_resets_is_held in tests/test_lock.py to confirm that cancelling a blocking acquire() task resets _is_held and allows subsequent acquisition calls on the same instance to succeed.

All 626 tests, ruff, mypy --strict, and 100% coverage on lock.py are passing. Pushed as commit 09b1fdb — ready for review/merge!

@allen0099

Copy link
Copy Markdown
Owner

Thanks, the cancellation fix and its test look good. I checked that the new test fails if except BaseException: is reverted to except Exception:. With live Redis and Memcached, the full suite passes locally: 790 tests, 100% coverage, mypy --strict clean.

ruff format --check flagged a trailing blank line at the end of tests/test_lock.py again, so I pushed 7cfe73f to remove it. Please git pull before any further push. Running uv run pre-commit install once in your clone runs ruff format on every commit and catches this automatically.

From my side this is ready. Please mark it ready for review.

@ShivanshShukla
ShivanshShukla marked this pull request as ready for review September 26, 2026 10:22
@allen0099
allen0099 merged commit 364fa04 into allen0099:master Sep 26, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a CacheLock helper on top of set_if_absent / delete_if_equals

2 participants