Skip to content

build: use ucontext coroutines instead of pthreads on Linux and arm64 - #3

Closed
jbaczuk-qualia wants to merge 2 commits into
async-resourcefrom
ucontext-arm64
Closed

jbaczuk-qualia wants to merge 2 commits into
async-resourcefrom
ucontext-arm64

Conversation

@jbaczuk-qualia

@jbaczuk-qualia jbaczuk-qualia commented Sep 15, 2026 •

Copy link
Copy Markdown

Summary

Production runs arm64 (a prod pod has node_modules/fibers/bin/linux-arm64-108-glibc, which only exists when npm compiled fibers from source: the published package ships linux-x64 prebuilts only). On arm64 binding.gyp selects CORO_PTHREAD, so every fiber is an OS thread and every Fiber.yield() / run() is a pthread condvar handoff between two threads. The same applies to the arm64 CI runners and to M-series dev containers.

This switches Linux to CORO_UCONTEXT on glibc (CORO_ASM on musl, upstream's conditions) and makes the arm/arm64 blocks select CORO_UCONTEXT explicitly. glibc implements swapcontext on aarch64; the "problems getting real fibers working on arm" comment predates arm64. Independent of the node 24 work in #2 (this is the same binding.gyp change, cherry-picked onto async-resource).

Numbers (node 18.16.1, arm64, same host)

Pure switch microbenchmark, 200k run()/yield() round trips:

Build yield -> caller p50 / p90 / p99
fibers 5.0.4 as compiled on arm64 today (CORO_PTHREAD) 12.7 / 15.3 / 22.5 µs
this branch (CORO_UCONTEXT) 0.71 / 0.88 / 2.2 µs

In global-deployment-center with @qualia/prom-client's fiber_yield_switch_seconds (qualialabs/qualia#56265), the pthread build measured 18 µs p50 / 577 µs p99 per switch during a manual click-through; the ucontext build on node 24 measured 2.0 µs / 48 µs on a similar workload. Beyond latency, pthread mode also costs one OS thread and one kernel context switch per fiber.

Verification

  • test/*.js on node 18.16.1 arm64: 18/19 with ucontext (pool.js and cleanup.js now pass) versus 16/19 with the pthread build; future-exception.js fails on both.
  • 200k-yield leak check (forced GC each round, long-lived fibers and fiber churn): heap slope < 1 B/yield, external memory flat, RSS plateaus. Same profile as the pthread build.
  • GDC on node 18 boots and serves with the ucontext binary swapped in (login, subs, methods over DDP).
  • Same host class as production (Graviton2 remote-dev box, node 18.16.1 on the nodejs-dev:f031bb860e image): suite 18/19 vs 16/19, switch p50 2.49 µs vs 14.0 µs, GC stress (--stress-incremental-marking --stress-compaction) 150 rounds clean on both backends, 200k-yield leak check flat on both. Production is Graviton4 (Neoverse V2), which only widens the gap.
  • Wasm code GC with a suspended fiber still FATALs on both backends; that is a V8 bug fixed separately in v8: report live wasm code from archived threads instead of FATAL (node 18 port) node#5 (node 18) and fix: reset the handle when a suspended fiber is garbage-collected on V8 >= 10.4 #4 (node 24), not something this change affects.
  • Not yet run: a Lucky suite on the arm64 CI runner with this build. That runner exercises the new backend for every service, so it is the gate before publishing.

Rollout

5.0.5 is already published (@meteor/blaze depends on ^5.0.5), so 5422b70 bumps this branch to 5.0.6. Publishing it makes every ^5.0.4 / ^5.0.5 consumer pick up the ucontext build on its next lockfile refresh; arm64 consumers compile from source at install time (no arm64 prebuilt is published), so the change takes effect on the next npm install in CI and in each service's Docker build. Intended, but it should follow the Lucky run above and ideally a canary on one Graviton4 pod.

🤖 Generated with Claude Code

fibers ships prebuilt binaries for linux-x64 only, so every arm64 host (production
Graviton nodes, the arm64 CI runners, M-series dev containers) compiles from source and
lands in the CORO_PTHREAD branch of binding.gyp, where each fiber is an OS thread and each
switch is a pthread condvar handoff. On the same node 18.16.1 arm64 host a
run()/yield() round trip costs 12.7 us p50 with pthreads and 0.71 us with swapcontext,
and the pthread build shows the same ~9x in Meteor's fiber_yield_switch_seconds
histogram. glibc implements swapcontext on aarch64; the "problems getting real fibers
working on arm" comment predates arm64.

Linux now picks CORO_UCONTEXT on glibc and CORO_ASM on musl (upstream's conditions),
and the arm/arm64 blocks select CORO_UCONTEXT explicitly. Test suite on node 18.16.1
arm64: 18/19 with ucontext (pool.js and cleanup.js now pass) versus 16/19 with pthreads;
future-exception.js fails on both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
5.0.5 is already published (@meteor/blaze depends on ^5.0.5), so the ucontext build needs a
new version before it can be released.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@jbaczuk-qualia

Copy link
Copy Markdown
Author

Closed 2026-09-21, kept for posterity. Switching arm64 production from CORO_PTHREAD to CORO_UCONTEXT is a separate decision from the node 24 upgrade, and for the upgrade we chose to keep pthread and change as little as possible (#5, qualialabs/node#6). The numbers here (4.8x p50 / 17x p99 per switch in GDC on the same node 18 image) remain valid if the backend question is reopened later; note that on node 24 ucontext also requires the V8 patches in qualialabs/node#4.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant