Skip to content

aarch64: master wedges inside TServer::Shutdown() — GracefulShutdownUnresponsiveSlave fails intermittently on the arm job #564

Description

@ohohoreilly

import_replication.test's GracefulShutdownUnresponsiveSlave fails intermittently on the aarch64 CI job. I first read this as load-sensitive flakiness; scoping the logs properly says otherwise. The master begins graceful shutdown and never finishes it.

Evidence

Latest occurrence: run 35281215846, master at dd770ac5.

The fixture gets all the way through its setup — its own markers are all present:

Suspending the slave
Slave state before interrupting the master: alive, state [T]
Interrupting the master

state [T] is stopped, so the slave was correctly frozen and the replication push was parked against it. The failure is in what happens after SIGINT.

Scoping to that fixture's own master log (the dump contains two fixtures' scratch dirs; this is the one that failed):

marker present
TServer::Shutdown() begin yes
TServer::Shutdown() draining [1] live client connections (#460) yes — and it is the last line
RunReplicationQueue shutting down (#461) no
TServer::Shutdown() complete no

So the master entered Shutdown(), logged the connection drain, and then produced nothing further. Not even the drain's own deadline warning, which is bounded at 10s and would have appeared. The fixture's first assertion after SIGINT is EXPECT_TRUE(master.Reap(seconds(60))) — it wedged and was reaped by force.

That is a hang in TServer::Shutdown(), not a slow test.

Rate

Across the 16 most recent ci.yml runs: aarch64 1 failure, amd64 0. The same fixture also failed on aarch64 on 2026-09-16; I re-ran that job to get a second sample, which replaced the logs, so I cannot show that one had the same signature — only that it was the same fixture on the same architecture. (Lesson noted: re-running a failed job to confirm a flake destroys the evidence you would need if it turns out not to be one.)

Nothing like it has been seen on amd64.

Why this should not be assumed a flake

The symptom is the same family as #554: a hang inside TServer::Shutdown() on aarch64 and nowhere else. That one turned out to be a fiber reading a cached thread pointer after migrating OS threads, which made a runner switch onto another thread's stack — fixed in #555 by routing the three fiber thread-locals through noipa accessors.

Three thread-locals in the same hazard class were deliberately left unaudited by that fix, because none was implicated and each is its own piece of work:

  • Disk::Util::TDiskController::TEvent::LocalEventPool
  • TLocalReadFileCache::Cache
  • Disk::TLocalWalkerCache::Cache

All three are read from fiber context and can be read after a switch. A stale read of any of them hands a fiber another thread's pool or cache. That is the first place I would look, though it is a hypothesis and not a diagnosis — the wedge could equally be an ordinary teardown-ordering bug that only arm's scheduling exposes.

What would make this diagnosable

Right now a failure yields log tails and nothing else, and the position in Shutdown() has to be inferred from which syslog lines are missing:

  1. Capture stacks when Reap times out. TChildServer::Reap failing is the moment to gdb -p the child, or at minimum dump /proc/<pid>/task/*/wchan and /proc/<pid>/task/*/syscall, before killing it. For orlyc hangs on aarch64 when compiling a package (embedded mem-sim server never returns) #554, field 8 of syscall — the user stack pointer — is what identified the bug when gdb was actively misleading on an LTO binary.
  2. The aarch64 job is continue-on-error, so this failure does not turn the run red and is invisible unless someone opens the job. That is the right setting for an unsupported-platform job and the wrong one now that v0.1.0 ships an arm64 image. Related: CI: guard the aarch64 orlyc path with a release build (debug does not reproduce arm bugs) #556.

Not urgent, but not nothing

Low frequency, and it is a graceful-shutdown path rather than a data path. But the project now publishes an arm64 image as part of a numbered release, and this is the second arm-only hang in TServer::Shutdown() in three days.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions