Repository navigation
[design] Run the sharded suite on zero Actions storage: cache and job outputs instead of artifacts, mode derived from visibility #82
Description
Activity
The gating measurement is largely answered, and my framing of it was wrong in a way that lowers the risk.
Measured — a
pull_requestrun in a private repo already writes cache. Existing entries in the fleet, scoped to the merge ref, from real PR runs:lab-sotashimozono/ParaLinearAlgebra.jl refs/pull/519/merge 126967 B refs/pull/589/merge 33617 B refs/pull/614/merge 144958 B (+2.5 MB entry) lab-sotashimozono/ComplexAnalysis.jl refs/pull/217/merge 134558 B refs/pull/219/merge 135149 B refs/pull/221/merge 142547 B (+2.5 MB entry)(Different workflow, same mechanism.) So read-write cache on a
pull_requestrun of a private repository is measured, not merely documented, and that is the path essentially every run in this fleet takes.And the fork question is narrower than the body claims. I wrote it as gating. It is not, because the design derives the mode from visibility:
- public repo → artifact path, unchanged. A fork PR of a public repo therefore never takes the cache path, so its cache permissions are irrelevant.
- private repo → cache path. A private repo can be forked by an org member, so a fork PR of a private repo is the one configuration where the question is live.
That is a narrow case, not a precondition for building. It should still be checked, and the honest fallback if it turns out read-only is small and local: on a fork PR, fall back to the artifact path — which is the same conditional the design already has for visibility, evaluated on one more term. Worst case, a fork PR of a private repo accrues a few KB of GB·hours, which is the situation today for every run.
Reordering: this is no longer a blocker before building, it is a case in the acceptance list.
Implementation constraint found while starting this, and it changes the shape of the change.
actions/cache/restorerestores exactly one key. There is nopattern:form, and no REST endpoint that downloads a cache by key (the cache API lists and deletes only).collecttherefore cannot mirror today's single- uses: actions/download-artifact@v8 with: { path: parts, pattern: ${{ inputs.artifact-prefix }}-* }
with a single restore. It needs one restore step per shard, and since Actions cannot generate steps dynamically, that means a bounded, literal list of them — plus a loud failure when
shardsexceeds the bound, because a silent cap here would report a complete run while skipping shards, which is the exact failure this workflow's completeness gate exists to catch.So the honest cost of the cache transport is higher than "swap the two steps": roughly 6 lines × the bound in
collect, again intimings, against ~10 lines today.A second payload turns out to be in scope, and it is not a convenience.
-lcovcarries the merged report fromcollect— which runs inside the reusable and therefore cannot seesecrets.CODECOV_TOKEN— out to the caller's own job, which can. That is a secret-boundary transport, and it is whyactions/upload-coverageexists at all (:89-98). Measured today, it is also the second of the two remaining failures:Coverage (upload): Unable to download artifact(s): Artifact not found for name: testshards-lcovNote the asymmetry with
collect, which is worth knowing generally:upload-coverageuses aname:download, which hard-fails when the artifact is absent, whilecollect'spattern:download succeeds with zero results. That is whycollectgets as far as the completeness gate and reportsunits observed = 0(correct — it refuses rather than passing blind), while the coverage job dies at the download.Measured state of the two jobs under a full quota, from run
32238414796— this is what the design has to fix, and nothing else incollectis broken:step result download-artifact(pattern)success — with zero artifacts Convert every shard's counters into one reportsuccess Publish the merged reportsuccess Every unit ran, exactly oncefailure — no shard reported how many units it observed — cannot verify the runeverything after skipped Before writing any of this into a General-registered package, the killing measurement is whether a cache saved in a matrix leg of
testis restorable by exact key incollectwithin the same run. The docs say yes for therefs/pull/.../mergescope; that is not a measurement.The killing measurement was run before writing any of this, and it kills the mechanism.
actions/cachecannot carry the payload in this repository.Measured on
QAtlasHub/TestShards.jlover four runs, with a throwaway workflow on a branch (branch and caches since deleted).Setup: a 3-leg matrix job writes a small file and calls
actions/cache/save@v4with keyprobe-out-<run_id>-<run_attempt>-<sid>; aneeds:-dependentcollectjob callsactions/cache/restore@v4with those exact keys.Save works, and the entries are real:
produce (s1): Cache saved with key: probe-out-32327386586-1-s1and the caches API lists them, on the right ref:
refs/heads/ci/probe-cache-transport probe-out-32327386586-1-s1 4500B 2026-08-20T03:12:19Z refs/heads/ci/probe-cache-transport probe-out-32327386586-1-s2 4464B 2026-08-20T03:12:20Z refs/heads/ci/probe-cache-transport probe-out-32327386586-1-s3 269B 2026-08-20T03:12:20ZRestore misses every one of them, with no error and no service warning:
Cache not found for input keys: probe-out-32327386586-1-s1Three controls, each ruling out an explanation:
control result rules out 45 s sleep between save and restore still a miss propagation delay a key from a previous run, confirmed present via the API 3 minutes earlier miss any same-run race a default-branch key written by this repo's own CI ( testshards;os=Linux;run_id=32210565474;run_attempt=1), listed as activemiss "non-default-ref caches are unreadable" a key never written miss (correct) — So restore does not fail for the payload, or the ref, or the timing: it does not restore anything in this repository, while save reports success and the API lists the entry as active. Root cause not established — the step emits a clean miss with correct inputs and no diagnostic.
Consequences
- The transport proposed in this issue is refuted. Not "expensive" — non-functional. Nothing in the design survives that depends on
shards → collectvia cache, which is the completeness evidence and the coverage counters, i.e. both of the two jobs that are currently failing. julia-actions/cachehas very likely never been restoring here either — same service, same repo. It would report success and re-do the work every run. This is measured forQAtlasHub/TestShards.jlonly; thelab-sotashimozonorepos are a different account and were not measured. Worth measuring before anyone relies oncache: true, and it is a plausible part of why#24foundcache: falsebetter on the persistent-depot runners.- The git-ref transport is back on the table, and it deserves a fair re-read, because two of the three reasons I rejected it in [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80 were themselves wrong:
"it would newly require— that is the status quo: thecontents: writeon jobs running arbitrary test code"testjob inherits the caller's grant, callers are required to grant it (:249-252), and:595already injectssecrets.GITHUB_TOKENinto the test process unconditionally.- "deleted refs leave unreachable objects" — true, but
ci-timingsis force-pushed every run by the same argument, so this is a size argument, not a principled one. - "a fork
pull_requestgets a read-only token" — this one stands, and is the only one that does. Under the visibility-derived mode it applies to public repos, which keep the artifact path anyway, so it constrains a narrow case rather than the design.
What is actually left
For
shards → collectwithout artifact storage, after this measurement:- git refs — works, measured end to end locally; needs the fork-PR fallback
- the Actions API, reading each shard's own job log from
collect—actions: read, no storage, no writes, works on fork PRs; fragile, since it parses a log cache— refuted above
labels → shardsvia job output is unaffected by any of this and remains strictly better than today's artifact (903 B, fork-safe, and it makes the history-divergence failure at:290-293structurally impossible).Reopening the transport question rather than the goal. The requirement — zero
Actions storage, all five outputs — is unchanged.- The transport proposed in this issue is refuted. Not "expensive" — non-functional. Nothing in the design survives that depends on
A transport that works, measured end to end: each shard emits its evidence into its own job log, and
collectreads it back through the Actions API.run=32328961422 completed/success produce jobs: 96305777760 96305777898 96305777914 recovered s3 from job 96305777760 (26 B) recovered s1 from job 96305777898 (26 B) recovered s2 from job 96305777914 (26 B) --- LOG transport recovered 3 of 3 --- OK-A: mid-run log read works; N->1 fan-in with NO storage and NO write tokenThe go/no-go was whether a finished job's log is readable from a later job of the same, still in-progress run. It is.
Why this is better than the artifact path it replaces, not merely equal
artifacts (today) cache (#82 as proposed) job logs Actions storagemetered — the whole problem none none works at all here yes no — refuted by measurement yes token needed runtime token runtime token actions: readonlyfork PR yes untested yes — actions: readis grantedresidue to clean up retention LRU none re-run isolation run-scoped (mixes attempts) key must carry run_attempt/attempts/{n}/jobs— exactIt needs no write permission at all, which also disposes of the
contents: writequestion that #80 got wrong in both directions.Four things that have to be right, each of which bit during the measurement
gh api .../actions/jobs/{id}/logsdoes not follow the redirect. The endpoint answers302to a short-lived Azure blob URL;gh apiexits 1 with no useful message.curl -sSLworks. Measured:status=302 redirect=https://productionresultssa11.blob.core.windows.net/actions-results/.../job-logs.txt?...- The runner echoes the step script into the log before running it, so a literal marker string appears twice — once as source, once as output.
head -1therefore reads back the unexpanded$(base64 -w0 ev.tsv)and the decode dies withbase64: invalid input. Take the last match and exclude$from the captured groups. - Use
/runs/{id}/attempts/{n}/jobs, not/runs/{id}/jobs, so a re-run reads its own attempt's jobs rather than merging attempts — the mirror of the flaw the cache design neededrun_attemptin the key to avoid, and which today's artifacts (run-scoped, not attempt-scoped) actually have. - Base64 the payload. A tab or a newline in a unit name would otherwise break the framing, and the marker has to be distinguishable from anything the suite itself prints.
Payload sizes, and why #74 becomes a prerequisite for the coverage half
One base64 line per shard:
payload raw base64 -out(shard-*.tsv,ran-*.tsv, timings, sections)7–9 KB ~10–12 KB coverage counters, today ~104 KB ~140 KB coverage counters, after #74/#75's counter_index~15 KB ~20 KB 12 KB on one line is unremarkable; 140 KB is asking for trouble. So shipping #75's workflow side stops being an optimisation and becomes the thing that makes the coverage payload fit this channel — the two changes converge rather than compete.
Also measured, and negative
actions/cachecannot carry this — see the previous comment. Save succeeds and the entry is listed as active; restore misses every key, including a default-branch entry written by this repo's own CI. Not a delay, not a scope rule.- The git-ref transport did not work inside Actions either: with
contents: writegranted andpersist-credentialsat its default, no ref reached the remote (git ls-remote origin 'refs/ci-run/*'was empty from the collector). Not debugged further, because the log transport needs no write token at all and is therefore strictly preferable. [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80's rejection of it stands, for better reasons than the ones it gave.
Next
Implement, in this order:
Coverage (upload)— tolerate a missing-lcov. It is aname:download, which hard-fails on absence, whilecollect'spattern:download succeeds with zero results. One conditional; it turns a red job into a reported degradation.collect's completeness evidence — over the log channel, keeping the check itself untouched.- Per-shard coverage artifacts ship 8 copies of the source tree — 85.7% of the payload is text
collectalready has checked out #74/Coverage counters can travel without the source they are printed against (#74) #75's wiring, then the coverage counters over the same channel. labels → shardsvia job output — unaffected by all of the above and still strictly better than its artifact.
The probe branch and its caches have been deleted.
Closing — keeping the current behaviour. The requirement was real and the transport question was answered; the answer is being left on the shelf rather than built.
What is settled, and does not need re-measuring:
question answer why artifacts stop Actions storageis accrued in GigabyteHours;CreateArtifactrefuses before any bytes are sent, and deleting does not un-accrue the periodactions/cacheas transportdoes not work here — save succeeds and lists active, restore misses every key, including a default-branch entry and with a 45 s gap. Four runs, three controls job logs as transport works — recovered 3 of 3in a live run; no storage, no residue,actions: readonly, fork-PR safegit refs as transport pushes did not reach the remote from inside Actions; not debugged, since the log channel needs no write token job outputs fine for labels → shards(903 B); cannot serveshards → collect— matrix legs do not merge outputs, last leg winsAnd three implementation facts that cost a run each to learn:
- A new action and its first reference cannot land in one PR (
uses:takes no expression; a relative reference in a reusable workflow resolves in the caller's checkout). gh api …/actions/jobs/{id}/logsanswers 302 and gh does not follow it.curl -sSLdoes.- The API names a reusable workflow's job with the caller's prefix —
ci / test / shard s5. A name filter therefore depends on what the caller named its job, and one that matches nothing reads zero logs and reports a complete-looking empty run.
The state this leaves:
Coverage (upload)and the completeness gate stay red on every private consumer until the billing period rolls over. The gate failing is correct — it refuses rather than passing blind — and that is the behaviour being kept.- A new action and its first reference cannot land in one PR (
Requirement: consume zero
Actions storage, and keep all five outputs. Not "fewer artifacts" — zero, because the meter this hits cannot be relieved by deleting anything.This supersedes the goal of #80 while rejecting its mechanism and its measurement. Read #80's body first; it records why the obvious decomposition does not work.
Why this is a hard requirement, not a preference
Actions storageis metered in GigabyteHours — an accrual over the billing period, not a snapshot of bytes held. From the billing usage report for the affected account:Three things follow, and each was got wrong before being measured:
0.5 GB × 730 h = 365 GB·h. August stands at 374.04. Crossed.net $0.00); storage has zero offset, so it is past the allowance and billable. With the default $0 spending limit, that is a refusal atCreateArtifact.So the block lifts on 2026-09-01 and not before, and it will recur, because nothing about the accrual model changes next month. Raising the spending limit would clear it today for $0.13; that is declined, which makes storage-free operation the design rather than a stopgap.
The transport that is not
Actions storageactions/cacheis a different resource, and this is measured rather than assumed — it is the fact the whole design rests on, so it goes first:Actions storagemetered in AugustActions Linux(Minutes),Actions storage(GigabyteHours) — no cache SKUIf cache were on that meter, 9.11 GB would accrue 219 GB·h per day and August's total could not be 374. It is not on that meter.
Two honest caveats:
ITensorPartitionFunctions.jl, 9.09 GB). It does not run sharded tests, so it is not affected by this design, but it is a monitoring item: if it crosses 10 GB the overage may land on the same meter this design exists to avoid.refs/pull/.../merge)" and later jobs of that run can restore it — which is precisely thetest → collectcase.What must NOT be weakened
#80 tried to delete the completeness check's transport by decomposing the claim. That fails, and the reasons are worth restating so this design does not retry it:
assignis total overkeys(timings), not over the units. Anything the history has not seen is placed by_owns(::Assigned, …)'sctx.unknown += 1— a mutable per-process counter in observation order. On a fresh consumerassignreturnsDict()and 100 % of ownership comes from that path.ctx.ranis appended iff_ownssaid so, from the samectx.assignment, in the same process.stealthere is no assignment at all (_owns(::Claimed, …)→_claim), so no local quantity exists to compare against.Conclusion: the evidence has to move. This design moves it through cache instead of through artifact storage. The check itself is untouched.
Direction by direction
timings.tsv,labels → shardsshard-*.tsv/ran-*.tsv/timings-*.tsv/sections-*.tsv/records-*.jsonl,shards → collect-out-*, 7–9 KB × N…-out-${run_id}-${run_attempt}-${sid}run_attemptin the key, so a partial re-run cannot merge two attempts' evidenceshards → collect-coverage-*, ~104 KB × Ncollectstill merges and uploads onceci-timingsbranch, fed from-out-*ci-timingsbranch, fed from cache-lcovartifact-prebuild-records30-day archiveWhy the job output is an improvement and not just a swap. Every shard receives the same string, so the failure this repo has already measured —
sharded-tests.yml:290-293, seven shards loading 23 timing rows and the eighth loading none after a transientgit ls-remote, "two units ran twice, two ran nowhere, and only the completeness gate noticed" — becomes structurally impossible rather than merely detected. 903 B against a ~1 MB output limit.needs.labels.outputs.timingsalready exists (:267,:482); this carries the content instead of a status.Why
shards → collectcannot use job outputs. GitHub does not merge outputs across matrix legs — the last leg to finish overwrites, order unspecified. That is the real reason this direction needs a store, and it is the argument #80 failed to make.The one thing that genuinely cannot survive. A downloadable merged lcov file requires somewhere to put a file. Without artifact storage the file is gone; the number is not. So: the merged percentage and the worst-covered files go into the step summary (free, no storage, readable in the UI), and the artifact stays available behind an input for anyone who wants the file and can afford it. What must not happen is a per-shard percentage — percentages do not average, only line sets union, and this repo already recorded the symptom: "reported 54.5 % for a suite that covered 94.8 %".
Derive the mode from visibility; do not ask the caller
Public repositories are not billed for Actions storage. Empirically:
QAtlasHub/TestShards.jlis public and its own CI — using this workflow with the then-unguarded uploads — ran green repeatedly inside the outage window, while every private consumer was red.So the switch should be derived from the caller's visibility, not exposed as an input:
Precedent in this fleet:
lab-sotashimozono/.github#27already derives the runner pool from the caller's visibility rather than asking every caller.This also disposes of the strongest objection to the whole enterprise: a fork pull request only happens on a public repository, where storage is free and the artifact path therefore keeps working untouched. No input default changes, and no external contributor sees different behaviour.
The gating measurement, to run BEFORE building anything
Documented to work, not measured here — and if it fails the design needs a different store, so it comes first:
The docs say a
pull_requestrun gets read-write cache in its ownrefs/pull/N/mergescope, and the 2026-06-26 read-only-cache changelog namespull_requestas keeping read-write. Neither statement is a measurement. One scratch fork PR settles it.Secondary, also unmeasured: the exact job-output size limit (secondary sources say 1 MB; 903 B makes the margin irrelevant, but anything larger would need it confirmed).
Acceptance
Every unit ran, exactly oncestill fails when it should, verified by the falsifiable fixture — two shards given different timing histories, each internally consistent — not by hand-editing one assignmentrun_attemptkey)shards: 1andsteal: trueboth still passdiagnose, defaulttrue) still reports peak concurrency and arrival spread — it is definitionally cross-shard and [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80's table omitted itRelation to #81 and #74
#81 (retention, and the write-token leak at
:595) and #74 (#75's workflow wiring) are recurrence prevention, not recovery — neither lifts this month's block, because neither un-accrues GB·hours. #81's item 1 is corrected accordingly. This issue is the one that makes a run survive the condition.