You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
MemcachedBackend uses pymemcache's HashClient with ignore_exc=False and the default retry_attempts=2, retry_timeout=1 and dead_timeout=60. When a server becomes unreachable, only some calls raise. For about a second after each failure, HashClient returns the command's default value (None or False) without raising. The backend then reads that value as a normal answer.
Timeline for a single unreachable server, calling get() every 100 ms:
0.0s raises ConnectionRefusedError
0.1s -> None (silent)
1.0s raises ConnectionRefusedError (retry)
1.1s -> None (silent)
2.0s raises ConnectionRefusedError (retry)
2.2s raises MemcacheError: All servers seem to be down right now (until the end of the 66 s trace)
What each method returns during a silent window, right after a failed call:
method
result
consequence
set
returns normally
the write is lost; set does not check pymemcache's return value
get
None
looks like a cache miss
set_if_absent
False
looks like the key is held by someone else
increment
0
a rate limiter built on it sees a fresh counter and lets the request through
Configure HashClient so that a failed server is never answered with a default. For example, retry_attempts=0 with a short dead_timeout, or a pymemcache client without the retry bookkeeping when there is one server. Recovery time after the server comes back needs checking.
Or detect the silent path in the backend: check set's return value, and treat None from gets / incr as an error where the protocol cannot return it for a reachable server.
Problem
MemcachedBackenduses pymemcache'sHashClientwithignore_exc=Falseand the defaultretry_attempts=2,retry_timeout=1anddead_timeout=60. When a server becomes unreachable, only some calls raise. For about a second after each failure,HashClientreturns the command's default value (NoneorFalse) without raising. The backend then reads that value as a normal answer.Timeline for a single unreachable server, calling
get()every 100 ms:What each method returns during a silent window, right after a failed call:
setsetdoes not check pymemcache's return valuegetNoneset_if_absentFalseincrement0get_and_deleteNonedelete_if_equals/expire_if_equalsTypeError: cannot unpack non-iterable NoneType objectgetsreturnedNoneinstead of a(value, cas)pairdelete_many(["a", "b"])2delete/clear_pathNone/0The Redis and memory backends never return a made-up answer for a failed call.
Repro
Possible directions
HashClientso that a failed server is never answered with a default. For example,retry_attempts=0with a shortdead_timeout, or a pymemcache client without the retry bookkeeping when there is one server. Recovery time after the server comes back needs checking.set's return value, and treatNonefromgets/incras an error where the protocol cannot return it for a reachable server.clear_pathin fix(memcached): let clear_path raise connection errors #196.Found while fixing #177.