A shard died on the default branch with
ERROR: LoadError: The following 1 direct dependency failed to precompile:
SystemError: opening file "…/.julia/compiled/v1.12/Markdown/AREjX_vfzSw.ji": No such file or directory
One shard read a cache file while another was replacing it. On a self-hosted pool the depot is persistent and shared per box, so N shards on the same box precompile into the same ~/.julia/compiled.
Why this is worth writing down rather than re-running past
The probability is a function of how many jobs precompile concurrently, and that number just changed by an order of magnitude. Before today, one repository ran one test job. Today thirty-odd repositories each run four to sixteen shards, on the same pool, and they will keep doing that on every push. The migration did not introduce the hazard; it moved it from "essentially never" to "expect it".
The failure is at least loud, and a re-run clears it. But a fleet where any run can fail for a reason unrelated to the code is a fleet where red checks stop being read — which is the thing several of this package's other guards exist to prevent.
What might fix it
julia-actions/cache already handles the hosted case; the self-hosted case is the one with a shared writable depot.
- A per-shard depot —
JULIA_DEPOT_PATH under the runner's workspace, with the shared one appended read-only. Costs precompilation time per shard, which is exactly what the shared depot was buying.
prebuild already exists and sidesteps it by precompiling once — but it is documented hosted-only, because Julia keys a package's compiled cache on its absolute source path and that path carries the runner's name on a self-hosted pool. Worth re-checking whether that reasoning still holds for DEPENDENCIES, which is where the race is: the failing file was Markdown, a stdlib, not the package under test.
That last point may be the cheap fix: dependencies do not change per commit and their caches could be built once and shared read-only, leaving only the package under test to be compiled per shard.
Status
Two repositories hit it on the same afternoon; both re-runs are pending as I write this. Filing before knowing whether they clear, because "the re-run was green" is not evidence the race is gone — it is evidence that a race is a race.
A shard died on the default branch with
One shard read a cache file while another was replacing it. On a self-hosted pool the depot is persistent and shared per box, so N shards on the same box precompile into the same
~/.julia/compiled.Why this is worth writing down rather than re-running past
The probability is a function of how many jobs precompile concurrently, and that number just changed by an order of magnitude. Before today, one repository ran one test job. Today thirty-odd repositories each run four to sixteen shards, on the same pool, and they will keep doing that on every push. The migration did not introduce the hazard; it moved it from "essentially never" to "expect it".
The failure is at least loud, and a re-run clears it. But a fleet where any run can fail for a reason unrelated to the code is a fleet where red checks stop being read — which is the thing several of this package's other guards exist to prevent.
What might fix it
julia-actions/cachealready handles the hosted case; the self-hosted case is the one with a shared writable depot.JULIA_DEPOT_PATHunder the runner's workspace, with the shared one appended read-only. Costs precompilation time per shard, which is exactly what the shared depot was buying.prebuildalready exists and sidesteps it by precompiling once — but it is documented hosted-only, because Julia keys a package's compiled cache on its absolute source path and that path carries the runner's name on a self-hosted pool. Worth re-checking whether that reasoning still holds for DEPENDENCIES, which is where the race is: the failing file wasMarkdown, a stdlib, not the package under test.That last point may be the cheap fix: dependencies do not change per commit and their caches could be built once and shared read-only, leaving only the package under test to be compiled per shard.
Status
Two repositories hit it on the same afternoon; both re-runs are pending as I write this. Filing before knowing whether they clear, because "the re-run was green" is not evidence the race is gone — it is evidence that a race is a race.