Skip to content

fix(memcached): raise while a server is unreachable instead of returning defaults - #216

Merged
allen0099 merged 1 commit into
masterfrom
fix/memcached-silent-defaults
Sep 26, 2026
Merged

allen0099 merged 1 commit into
masterfrom
fix/memcached-silent-defaults

Conversation

@allen0099

Copy link
Copy Markdown
Owner

Closes #197

Problem

MemcachedBackend built its HashClient with pymemcache's default retry settings (retry_attempts=2, retry_timeout=1, dead_timeout=60). After a call to a server failed, every call to that server returned the command's default value without raising, until the retry was due. On a single-server backend, one failed get() made the next calls return:

method result before
get / get_and_delete None (looks like a miss / already consumed)
set returns normally, write lost
set_if_absent False (looks like the key is taken)
increment 0 (a rate limiter sees a fresh counter)
delete_if_equals / expire_if_equals TypeError from unpacking None
delete_many(["a", "b"]) 2
delete / clear_path None / 0

After two more failures the server was marked dead for 60 seconds.

Change

The client is now built with retry_attempts=0 and dead_timeout=1. With ignore_exc=False (unchanged), this leaves no path in HashClient that returns a default:

  • The first failure raises the connection error and takes the server out of rotation at once.
  • Calls in the next second raise MemcacheError("All servers seem to be down right now") for a single server.
  • After one second the server is tried again. A server that is back is used again after about one second, instead of up to about 60.

With several servers, a failed server's keys go to the remaining servers until it answers again. That is HashClient's normal failover; before this change it started after the third failure instead of the first.

Docs (en and zh-TW BACKENDS.md) describe the behaviour, and the CHANGELOG has a Fixed entry.

Verification

  • Repro from the issue, all ten methods, before and after: every method now raises on the call right after a failure.
  • Recovery against a live memcached that was stopped and started again: the backend raised while it was down and wrote successfully 1.0 s after it was back.
  • New tests in tests/backends/test_memcached.py. They need no server:
    • test_memcached_keeps_raising_after_a_connection_failure, parametrized over the ten methods;
    • test_memcached_tries_a_failed_server_again_after_the_dead_timeout.
  • Mutation checks:
    • dropping retry_attempts=0 fails the ten parametrized cases and nothing else;
    • dropping dead_timeout fails only the dead-timeout test.
  • Ruff, mypy --strict, the full suite against live Redis and Memcached (893 passed), and zensical build --strict for both languages.

Note

This does not touch get_and_delete's logic, so it does not conflict with the CAS change planned in #175.

…ing defaults

With pymemcache's default retries, HashClient answered every call to a
server that had just failed with the command's default until the retry
was due: get() looked like a miss, set() dropped the write, increment()
returned 0 and delete_if_equals() raised TypeError on the missing CAS
pair. The client now uses retry_attempts=0, so a failed server leaves
rotation at once and calls raise, and dead_timeout=1, so it is tried
again after one second rather than 60.

Closes #197
@allen0099 allen0099 added this to the 0.3.8 milestone Sep 26, 2026
@allen0099 allen0099 added bug Something isn't working backends Cache backends and their atomic primitives labels Sep 26, 2026
@allen0099
allen0099 merged commit e976a58 into master Sep 26, 2026
11 checks passed
@allen0099
allen0099 deleted the fix/memcached-silent-defaults branch September 26, 2026 16:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backends Cache backends and their atomic primitives bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Memcached backend returns made-up results while a server is unreachable

1 participant