Skip to content

ci(julia-testshards): default cache: false — this reusable's runner has a persistent depot - #24

Merged
sotashimozono merged 1 commit into
mainfrom
cache-false-selfhosted
Aug 5, 2026
Merged

sotashimozono merged 1 commit into
mainfrom
cache-false-selfhosted

Conversation

@sotashimozono

Copy link
Copy Markdown
Contributor

The other half of QAtlasHub/TestShards.jl#62, whose fix (a cache input) is QAtlasHub/TestShards.jl#71.

TestShards defaults cache: true, which is right for its callers — a hosted runner's depot is
gone at the end of the job. This reusable defaults runner to '["self-hosted","rosina"]',
whose depot persists per box, so the same default is exactly backwards here.

Measured

Post Run julia-actions/cache@v3 per shard 44–182 s
actual testing per shard 58–268 s
runner time per 8-shard run, for a depot already on the box ≈620 s

And it does not merely cost. Three jobs on 2026-08-03 whose test step succeeded went red or never
finished in that post-step — one held a runner for 114 minutes. A hung post-step occupies a
self-hosted runner indefinitely, so it shrinks the pool for everyone and the symptom downstream is
"CI is queued", which reads as ordinary contention rather than as a fault.

Seen again on ParaLinearAlgebra.jl run 30961107784: six shards on rosina spent 4–5 minutes each in
that post-step and finished; two on panza were still in it 16 minutes later with every test green.

Why an input rather than a hardcoded false

The caller is what knows whether its depot survives the job. A caller that overrides runner to a
hosted one can set cache: true and get the behaviour that is right for it.

Merge after QAtlasHub/TestShards.jl#71, which is where the input comes from.

… has a persistent depot

TestShards' `cache` input defaults to `true`, which is right for its own callers: a hosted runner's
depot is gone at the end of the job. This reusable defaults `runner` to `'["self-hosted","rosina"]'`,
whose depot persists per box, so here the same default is exactly backwards.

Measured (QAtlasHub/TestShards.jl#62): 44–182 s of `Post Run julia-actions/cache@v3` per shard
against 58–268 s of actual testing, ≈620 s of runner time per run, for a depot already on the box.
And it does not merely cost — three jobs on 2026-08-03 whose test step SUCCEEDED went red or never
finished in that post-step, one holding a runner for 114 minutes, which reads downstream as ordinary
queueing rather than as a fault.

Passed through rather than hardcoded, so a caller that overrides `runner` to a hosted one can turn
it back on — the caller is what knows whether its depot survives the job.

Refs QAtlasHub/TestShards.jl#62, #71.
@sotashimozono
sotashimozono merged commit fb981ff into main Aug 5, 2026
1 check passed
@sotashimozono
sotashimozono deleted the cache-false-selfhosted branch August 5, 2026 00:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant