Problem
MemcachedBackend.increment(key, ttl=1) can raise CacheXError("Counter vanished between ADD and INCR") on a key that no other caller touches.
This happened in the coverage job on PR #314 (run 36324448236). The failing test was tests/backends/test_memcached.py::test_memcached_increment_honors_ttl, on its first call increment("window", ttl=1). PR #314 does not touch the Memcached backend. This was the only failure of that workflow in its last 40 runs, so it is rare.
Cause
Memcached keeps time at one-second resolution. An item stored with exptime 1 at internal time T gets exptime = T + 1, and it counts as expired once the clock reaches T + 1. The clock ticks once a second at a moment unrelated to the write, so an item with ttl=1 can live anywhere from about 0 to 1 second.
_increment runs INCR (a miss), then ADD with that exptime, then INCR again. If the clock ticks between the ADD and the second INCR, the counter has already expired and increment raises. The same applies to any ttl whose expiry falls between the two commands, but with ttl=1 it can happen on a fresh key.
Proposal
When the second INCR misses, retry ADD + INCR a bounded number of times, the way the CAS paths retry (_CAS_MAX_RETRIES), and raise only after that. A counter that expired inside its own window really is a new window, so starting again at delta is correct.
Add a unit test with a stubbed client whose first INCR after ADD returns None.
Also consider whether the live test should use ttl=2 with a longer sleep, so that it stops depending on the tick.
Problem
MemcachedBackend.increment(key, ttl=1)can raiseCacheXError("Counter vanished between ADD and INCR")on a key that no other caller touches.This happened in the coverage job on PR #314 (run 36324448236). The failing test was
tests/backends/test_memcached.py::test_memcached_increment_honors_ttl, on its first callincrement("window", ttl=1). PR #314 does not touch the Memcached backend. This was the only failure of that workflow in its last 40 runs, so it is rare.Cause
Memcached keeps time at one-second resolution. An item stored with exptime
1at internal timeTgetsexptime = T + 1, and it counts as expired once the clock reachesT + 1. The clock ticks once a second at a moment unrelated to the write, so an item withttl=1can live anywhere from about 0 to 1 second._incrementrunsINCR(a miss), thenADDwith that exptime, thenINCRagain. If the clock ticks between theADDand the secondINCR, the counter has already expired andincrementraises. The same applies to any ttl whose expiry falls between the two commands, but withttl=1it can happen on a fresh key.Proposal
When the second
INCRmisses, retryADD+INCRa bounded number of times, the way the CAS paths retry (_CAS_MAX_RETRIES), and raise only after that. A counter that expired inside its own window really is a new window, so starting again atdeltais correct.Add a unit test with a stubbed client whose first
INCRafterADDreturnsNone.Also consider whether the live test should use
ttl=2with a longer sleep, so that it stops depending on the tick.