You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
import_replication.test's GracefulShutdownUnresponsiveSlave fails intermittently on the aarch64 CI job. I first read this as load-sensitive flakiness; scoping the logs properly says otherwise. The master begins graceful shutdown and never finishes it.
Evidence
Latest occurrence: run 35281215846, master at dd770ac5.
The fixture gets all the way through its setup — its own markers are all present:
Suspending the slave
Slave state before interrupting the master: alive, state [T]
Interrupting the master
state [T] is stopped, so the slave was correctly frozen and the replication push was parked against it. The failure is in what happens after SIGINT.
Scoping to that fixture's own master log (the dump contains two fixtures' scratch dirs; this is the one that failed):
marker
present
TServer::Shutdown() begin
yes
TServer::Shutdown() draining [1] live client connections (#460)
yes — and it is the last line
RunReplicationQueue shutting down (#461)
no
TServer::Shutdown() complete
no
So the master entered Shutdown(), logged the connection drain, and then produced nothing further. Not even the drain's own deadline warning, which is bounded at 10s and would have appeared. The fixture's first assertion after SIGINT is EXPECT_TRUE(master.Reap(seconds(60))) — it wedged and was reaped by force.
That is a hang in TServer::Shutdown(), not a slow test.
Rate
Across the 16 most recent ci.yml runs: aarch64 1 failure, amd64 0. The same fixture also failed on aarch64 on 2026-09-16; I re-ran that job to get a second sample, which replaced the logs, so I cannot show that one had the same signature — only that it was the same fixture on the same architecture. (Lesson noted: re-running a failed job to confirm a flake destroys the evidence you would need if it turns out not to be one.)
Nothing like it has been seen on amd64.
Why this should not be assumed a flake
The symptom is the same family as #554: a hang inside TServer::Shutdown() on aarch64 and nowhere else. That one turned out to be a fiber reading a cached thread pointer after migrating OS threads, which made a runner switch onto another thread's stack — fixed in #555 by routing the three fiber thread-locals through noipa accessors.
Three thread-locals in the same hazard class were deliberately left unaudited by that fix, because none was implicated and each is its own piece of work:
All three are read from fiber context and can be read after a switch. A stale read of any of them hands a fiber another thread's pool or cache. That is the first place I would look, though it is a hypothesis and not a diagnosis — the wedge could equally be an ordinary teardown-ordering bug that only arm's scheduling exposes.
What would make this diagnosable
Right now a failure yields log tails and nothing else, and the position in Shutdown() has to be inferred from which syslog lines are missing:
Capture stacks when Reap times out.TChildServer::Reap failing is the moment to gdb -p the child, or at minimum dump /proc/<pid>/task/*/wchan and /proc/<pid>/task/*/syscall, before killing it. For orlyc hangs on aarch64 when compiling a package (embedded mem-sim server never returns) #554, field 8 of syscall — the user stack pointer — is what identified the bug when gdb was actively misleading on an LTO binary.
Low frequency, and it is a graceful-shutdown path rather than a data path. But the project now publishes an arm64 image as part of a numbered release, and this is the second arm-only hang in TServer::Shutdown() in three days.
import_replication.test'sGracefulShutdownUnresponsiveSlavefails intermittently on the aarch64 CI job. I first read this as load-sensitive flakiness; scoping the logs properly says otherwise. The master begins graceful shutdown and never finishes it.Evidence
Latest occurrence: run 35281215846, master at
dd770ac5.The fixture gets all the way through its setup — its own markers are all present:
state [T]is stopped, so the slave was correctly frozen and the replication push was parked against it. The failure is in what happens after SIGINT.Scoping to that fixture's own master log (the dump contains two fixtures' scratch dirs; this is the one that failed):
TServer::Shutdown() beginTServer::Shutdown() draining [1] live client connections (#460)RunReplicationQueue shutting down (#461)TServer::Shutdown() completeSo the master entered
Shutdown(), logged the connection drain, and then produced nothing further. Not even the drain's own deadline warning, which is bounded at 10s and would have appeared. The fixture's first assertion after SIGINT isEXPECT_TRUE(master.Reap(seconds(60)))— it wedged and was reaped by force.That is a hang in
TServer::Shutdown(), not a slow test.Rate
Across the 16 most recent
ci.ymlruns: aarch64 1 failure, amd64 0. The same fixture also failed on aarch64 on 2026-09-16; I re-ran that job to get a second sample, which replaced the logs, so I cannot show that one had the same signature — only that it was the same fixture on the same architecture. (Lesson noted: re-running a failed job to confirm a flake destroys the evidence you would need if it turns out not to be one.)Nothing like it has been seen on amd64.
Why this should not be assumed a flake
The symptom is the same family as #554: a hang inside
TServer::Shutdown()on aarch64 and nowhere else. That one turned out to be a fiber reading a cached thread pointer after migrating OS threads, which made a runner switch onto another thread's stack — fixed in #555 by routing the three fiber thread-locals throughnoipaaccessors.Three thread-locals in the same hazard class were deliberately left unaudited by that fix, because none was implicated and each is its own piece of work:
Disk::Util::TDiskController::TEvent::LocalEventPoolTLocalReadFileCache::CacheDisk::TLocalWalkerCache::CacheAll three are read from fiber context and can be read after a switch. A stale read of any of them hands a fiber another thread's pool or cache. That is the first place I would look, though it is a hypothesis and not a diagnosis — the wedge could equally be an ordinary teardown-ordering bug that only arm's scheduling exposes.
What would make this diagnosable
Right now a failure yields log tails and nothing else, and the position in
Shutdown()has to be inferred from which syslog lines are missing:Reaptimes out.TChildServer::Reapfailing is the moment togdb -pthe child, or at minimum dump/proc/<pid>/task/*/wchanand/proc/<pid>/task/*/syscall, before killing it. For orlyc hangs on aarch64 when compiling a package (embedded mem-sim server never returns) #554, field 8 ofsyscall— the user stack pointer — is what identified the bug when gdb was actively misleading on an LTO binary.continue-on-error, so this failure does not turn the run red and is invisible unless someone opens the job. That is the right setting for an unsupported-platform job and the wrong one now that v0.1.0 ships an arm64 image. Related: CI: guard the aarch64 orlyc path with a release build (debug does not reproduce arm bugs) #556.Not urgent, but not nothing
Low frequency, and it is a graceful-shutdown path rather than a data path. But the project now publishes an arm64 image as part of a numbered release, and this is the second arm-only hang in
TServer::Shutdown()in three days.