Skip to content

Memcached backend returns made-up results while a server is unreachable #197

Description

@allen0099

Problem

MemcachedBackend uses pymemcache's HashClient with ignore_exc=False and the default retry_attempts=2, retry_timeout=1 and dead_timeout=60. When a server becomes unreachable, only some calls raise. For about a second after each failure, HashClient returns the command's default value (None or False) without raising. The backend then reads that value as a normal answer.

Timeline for a single unreachable server, calling get() every 100 ms:

 0.0s raises ConnectionRefusedError
 0.1s -> None            (silent)
 1.0s raises ConnectionRefusedError   (retry)
 1.1s -> None            (silent)
 2.0s raises ConnectionRefusedError   (retry)
 2.2s raises MemcacheError: All servers seem to be down right now   (until the end of the 66 s trace)

What each method returns during a silent window, right after a failed call:

method result consequence
set returns normally the write is lost; set does not check pymemcache's return value
get None looks like a cache miss
set_if_absent False looks like the key is held by someone else
increment 0 a rate limiter built on it sees a fresh counter and lets the request through
get_and_delete None looks like the value was already consumed
delete_if_equals / expire_if_equals TypeError: cannot unpack non-iterable NoneType object gets returned None instead of a (value, cas) pair
delete_many(["a", "b"]) 2 reports two deletions (see also #174)
delete / clear_path None / 0 looks like nothing was there

The Redis and memory backends never return a made-up answer for a failed call.

Repro

import asyncio, socket
from fastapi_cachex.backends import MemcachedBackend

s = socket.socket(); s.bind(("127.0.0.1", 0)); port = s.getsockname()[1]; s.close()
backend = MemcachedBackend(servers=[f"127.0.0.1:{port}"])

async def main():
    try:
        await backend.get("k")  # raises ConnectionRefusedError
    except OSError:
        pass
    print(await backend.increment("hits"))  # 0, no error

asyncio.run(main())

Possible directions

  • Configure HashClient so that a failed server is never answered with a default. For example, retry_attempts=0 with a short dead_timeout, or a pymemcache client without the retry bookkeeping when there is one server. Recovery time after the server comes back needs checking.
  • Or detect the silent path in the backend: check set's return value, and treat None from gets / incr as an error where the protocol cannot return it for a reachable server.
  • Either way, add a test per method against an unreachable server, like the one added for clear_path in fix(memcached): let clear_path raise connection errors #196.

Found while fixing #177.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    backendsCache backends and their atomic primitivesbugSomething isn't working

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions