Skip to content

[design] Run the sharded suite on zero Actions storage: cache and job outputs instead of artifacts, mode derived from visibility #82

Description

@sotashimozono

Requirement: consume zero Actions storage, and keep all five outputs. Not "fewer artifacts" — zero, because the meter this hits cannot be relieved by deleting anything.

This supersedes the goal of #80 while rejecting its mechanism and its measurement. Read #80's body first; it records why the obvious decomposition does not work.

Why this is a hard requirement, not a preference

Actions storage is metered in GigabyteHours — an accrual over the billing period, not a snapshot of bytes held. From the billing usage report for the affected account:

month quantity unit gross discount net
2026-07 2.5618 GigabyteHours $0.0009 — $0.0009
2026-08 374.0412 GigabyteHours $0.1257 $0.00 $0.1257
(compare) Actions Linux 4160 Minutes $24.96 $24.96 $0.00

Three things follow, and each was got wrong before being measured:

  1. The allowance is an accrual. A 500 MB Free-plan private allowance over a month is 0.5 GB × 730 h = 365 GB·h. August stands at 374.04. Crossed.
  2. Deleting does not help. Accrued GB·hours are not refunded. Live bytes are only 90.7 MB across 25 private repos, and uploads still fail — which is exactly the observation that no bytes-based model explains. The "usage is recalculated every 6–12 hours" message refers to the reported figure, not to the accrual resetting.
  3. The discount column proves the diagnosis. Minutes are fully offset by the included allowance (net $0.00); storage has zero offset, so it is past the allowance and billable. With the default $0 spending limit, that is a refusal at CreateArtifact.

So the block lifts on 2026-09-01 and not before, and it will recur, because nothing about the accrual model changes next month. Raising the spending limit would clear it today for $0.13; that is declined, which makes storage-free operation the design rather than a stopgap.

The transport that is not Actions storage

actions/cache is a different resource, and this is measured rather than assumed — it is the fact the whole design rests on, so it goes first:

cache held across the account 9.11 GB, 27 entries
Actions storage metered in August 374 GB·h ≈ 0.51 GB average
SKUs present in the billing report Actions Linux (Minutes), Actions storage (GigabyteHours) — no cache SKU

If cache were on that meter, 9.11 GB would accrue 219 GB·h per day and August's total could not be 374. It is not on that meter.

Two honest caveats:

  • Cache is free only to 10 GB per repository; past that, "any usage beyond 10 GB is billed to your account". Our numbers: every repo running sharded tests holds ≤ 0.01 GB, the payload is ~64 KB per run, and GitHub evicts entries unused for 7 days — so steady state is tens of MB against 10 GB.
  • One repo in the fleet is at 91 % of that threshold (ITensorPartitionFunctions.jl, 9.09 GB). It does not run sharded tests, so it is not affected by this design, but it is a monitoring item: if it crosses 10 GB the overage may land on the same meter this design exists to avoid.
  • Same-run restore is documented: "When a cache is created by a workflow run triggered on a pull request, the cache is created for the merge ref (refs/pull/.../merge)" and later jobs of that run can restore it — which is precisely the test → collect case.

What must NOT be weakened

#80 tried to delete the completeness check's transport by decomposing the claim. That fails, and the reasons are worth restating so this design does not retry it:

  • assign is total over keys(timings), not over the units. Anything the history has not seen is placed by _owns(::Assigned, …)'s ctx.unknown += 1 — a mutable per-process counter in observation order. On a fresh consumer assign returns Dict() and 100 % of ownership comes from that path.
  • "Each shard ran its assignment" is a tautology: ctx.ran is appended iff _owns said so, from the same ctx.assignment, in the same process.
  • Under steal there is no assignment at all (_owns(::Claimed, …) → _claim), so no local quantity exists to compare against.

Conclusion: the evidence has to move. This design moves it through cache instead of through artifact storage. The check itself is untouched.

Direction by direction

payload today proposed note
timings.tsv, labels → shards artifact, 903 B job output (base64) strictly better — see below
shard-*.tsv/ran-*.tsv/timings-*.tsv/sections-*.tsv/records-*.jsonl, shards → collect -out-*, 7–9 KB × N cache, key …-out-${run_id}-${run_attempt}-${sid} run_attempt in the key, so a partial re-run cannot merge two attempts' evidence
coverage counters, shards → collect -coverage-*, ~104 KB × N cache, same scheme collect still merges and uploads once
timing history, across runs ci-timings branch, fed from -out-* ci-timings branch, fed from cache the branch channel is unchanged
merged lcov, for a tokenless user -lcov artifact step summary + artifact behind an input the only genuine loss; see below
-prebuild artifact cache it is already cache-shaped
-records 30-day archive artifact behind an input, off in this mode purely forensic; 19 % of live bytes

Why the job output is an improvement and not just a swap. Every shard receives the same string, so the failure this repo has already measured — sharded-tests.yml:290-293, seven shards loading 23 timing rows and the eighth loading none after a transient git ls-remote, "two units ran twice, two ran nowhere, and only the completeness gate noticed" — becomes structurally impossible rather than merely detected. 903 B against a ~1 MB output limit. needs.labels.outputs.timings already exists (:267, :482); this carries the content instead of a status.

Why shards → collect cannot use job outputs. GitHub does not merge outputs across matrix legs — the last leg to finish overwrites, order unspecified. That is the real reason this direction needs a store, and it is the argument #80 failed to make.

The one thing that genuinely cannot survive. A downloadable merged lcov file requires somewhere to put a file. Without artifact storage the file is gone; the number is not. So: the merged percentage and the worst-covered files go into the step summary (free, no storage, readable in the UI), and the artifact stays available behind an input for anyone who wants the file and can afford it. What must not happen is a per-shard percentage — percentages do not average, only line sets union, and this repo already recorded the symptom: "reported 54.5 % for a suite that covered 94.8 %".

Derive the mode from visibility; do not ask the caller

Public repositories are not billed for Actions storage. Empirically: QAtlasHub/TestShards.jl is public and its own CI — using this workflow with the then-unguarded uploads — ran green repeatedly inside the outage window, while every private consumer was red.

So the switch should be derived from the caller's visibility, not exposed as an input:

  • private → cache + job outputs, zero artifacts
  • public → artifacts exactly as today, unchanged

Precedent in this fleet: lab-sotashimozono/.github#27 already derives the runner pool from the caller's visibility rather than asking every caller.

This also disposes of the strongest objection to the whole enterprise: a fork pull request only happens on a public repository, where storage is free and the artifact path therefore keeps working untouched. No input default changes, and no external contributor sees different behaviour.

The gating measurement, to run BEFORE building anything

Documented to work, not measured here — and if it fails the design needs a different store, so it comes first:

On a pull_request from a fork of a public repository, can a test job actions/cache save, and can collect restore it by exact key within the same run?

The docs say a pull_request run gets read-write cache in its own refs/pull/N/merge scope, and the 2026-06-26 read-only-cache changelog names pull_request as keeping read-write. Neither statement is a measurement. One scratch fork PR settles it.

Secondary, also unmeasured: the exact job-output size limit (secondary sources say 1 MB; 903 B makes the margin irrelevant, but anything larger would need it confirmed).

Acceptance

  • a run of a private consumer creates zero artifacts and still: runs all shards, gates completeness, reports merged coverage, and records timing history
  • Every unit ran, exactly once still fails when it should, verified by the falsifiable fixture — two shards given different timing histories, each internally consistent — not by hand-editing one assignment
  • every shard reports the same history row count, on every run
  • a partial re-run of one shard does not merge two attempts' evidence (the run_attempt key)
  • a public consumer's behaviour is byte-identical to today, including on a fork PR
  • shards: 1 and steal: true both still pass
  • the diagnosis (diagnose, default true) still reports peak concurrency and arrival spread — it is definitionally cross-shard and [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80's table omitted it
  • a tokenless user gets a coverage number, and never a per-shard percentage

Relation to #81 and #74

#81 (retention, and the write-token leak at :595) and #74 (#75's workflow wiring) are recurrence prevention, not recovery — neither lifts this month's block, because neither un-accrues GB·hours. #81's item 1 is corrected accordingly. This issue is the one that makes a run survive the condition.

Activity

  1. sotashimozono commented on Aug 19, 2026

    @sotashimozono
    MemberAuthor

    The gating measurement is largely answered, and my framing of it was wrong in a way that lowers the risk.

    Measured — a pull_request run in a private repo already writes cache. Existing entries in the fleet, scoped to the merge ref, from real PR runs:

    lab-sotashimozono/ParaLinearAlgebra.jl   refs/pull/519/merge   126967 B
                                             refs/pull/589/merge    33617 B
                                             refs/pull/614/merge   144958 B  (+2.5 MB entry)
    lab-sotashimozono/ComplexAnalysis.jl     refs/pull/217/merge   134558 B
                                             refs/pull/219/merge   135149 B
                                             refs/pull/221/merge   142547 B  (+2.5 MB entry)
    

    (Different workflow, same mechanism.) So read-write cache on a pull_request run of a private repository is measured, not merely documented, and that is the path essentially every run in this fleet takes.

    And the fork question is narrower than the body claims. I wrote it as gating. It is not, because the design derives the mode from visibility:

    • public repo → artifact path, unchanged. A fork PR of a public repo therefore never takes the cache path, so its cache permissions are irrelevant.
    • private repo → cache path. A private repo can be forked by an org member, so a fork PR of a private repo is the one configuration where the question is live.

    That is a narrow case, not a precondition for building. It should still be checked, and the honest fallback if it turns out read-only is small and local: on a fork PR, fall back to the artifact path — which is the same conditional the design already has for visibility, evaluated on one more term. Worst case, a fork PR of a private repo accrues a few KB of GB·hours, which is the situation today for every run.

    Reordering: this is no longer a blocker before building, it is a case in the acceptance list.

  2. sotashimozono commented on Aug 20, 2026

    @sotashimozono
    MemberAuthor

    Implementation constraint found while starting this, and it changes the shape of the change.

    actions/cache/restore restores exactly one key. There is no pattern: form, and no REST endpoint that downloads a cache by key (the cache API lists and deletes only). collect therefore cannot mirror today's single

    - uses: actions/download-artifact@v8
      with: { path: parts, pattern: ${{ inputs.artifact-prefix }}-* }

    with a single restore. It needs one restore step per shard, and since Actions cannot generate steps dynamically, that means a bounded, literal list of them — plus a loud failure when shards exceeds the bound, because a silent cap here would report a complete run while skipping shards, which is the exact failure this workflow's completeness gate exists to catch.

    So the honest cost of the cache transport is higher than "swap the two steps": roughly 6 lines × the bound in collect, again in timings, against ~10 lines today.

    A second payload turns out to be in scope, and it is not a convenience. -lcov carries the merged report from collect — which runs inside the reusable and therefore cannot see secrets.CODECOV_TOKEN — out to the caller's own job, which can. That is a secret-boundary transport, and it is why actions/upload-coverage exists at all (:89-98). Measured today, it is also the second of the two remaining failures:

    Coverage (upload):  Unable to download artifact(s): Artifact not found for name: testshards-lcov
    

    Note the asymmetry with collect, which is worth knowing generally: upload-coverage uses a name: download, which hard-fails when the artifact is absent, while collect's pattern: download succeeds with zero results. That is why collect gets as far as the completeness gate and reports units observed = 0 (correct — it refuses rather than passing blind), while the coverage job dies at the download.

    Measured state of the two jobs under a full quota, from run 32238414796 — this is what the design has to fix, and nothing else in collect is broken:

    step result
    download-artifact (pattern) success — with zero artifacts
    Convert every shard's counters into one report success
    Publish the merged report success
    Every unit ran, exactly once failure — no shard reported how many units it observed — cannot verify the run
    everything after skipped

    Before writing any of this into a General-registered package, the killing measurement is whether a cache saved in a matrix leg of test is restorable by exact key in collect within the same run. The docs say yes for the refs/pull/.../merge scope; that is not a measurement.

  3. sotashimozono commented on Aug 20, 2026

    @sotashimozono
    MemberAuthor

    The killing measurement was run before writing any of this, and it kills the mechanism. actions/cache cannot carry the payload in this repository.

    Measured on QAtlasHub/TestShards.jl over four runs, with a throwaway workflow on a branch (branch and caches since deleted).

    Setup: a 3-leg matrix job writes a small file and calls actions/cache/save@v4 with key probe-out-<run_id>-<run_attempt>-<sid>; a needs:-dependent collect job calls actions/cache/restore@v4 with those exact keys.

    Save works, and the entries are real:

    produce (s1): Cache saved with key: probe-out-32327386586-1-s1
    

    and the caches API lists them, on the right ref:

    refs/heads/ci/probe-cache-transport   probe-out-32327386586-1-s1   4500B   2026-08-20T03:12:19Z
    refs/heads/ci/probe-cache-transport   probe-out-32327386586-1-s2   4464B   2026-08-20T03:12:20Z
    refs/heads/ci/probe-cache-transport   probe-out-32327386586-1-s3    269B   2026-08-20T03:12:20Z
    

    Restore misses every one of them, with no error and no service warning:

    Cache not found for input keys: probe-out-32327386586-1-s1
    

    Three controls, each ruling out an explanation:

    control result rules out
    45 s sleep between save and restore still a miss propagation delay
    a key from a previous run, confirmed present via the API 3 minutes earlier miss any same-run race
    a default-branch key written by this repo's own CI (testshards;os=Linux;run_id=32210565474;run_attempt=1), listed as active miss "non-default-ref caches are unreadable"
    a key never written miss (correct) —

    So restore does not fail for the payload, or the ref, or the timing: it does not restore anything in this repository, while save reports success and the API lists the entry as active. Root cause not established — the step emits a clean miss with correct inputs and no diagnostic.

    Consequences

    1. The transport proposed in this issue is refuted. Not "expensive" — non-functional. Nothing in the design survives that depends on shards → collect via cache, which is the completeness evidence and the coverage counters, i.e. both of the two jobs that are currently failing.
    2. julia-actions/cache has very likely never been restoring here either — same service, same repo. It would report success and re-do the work every run. This is measured for QAtlasHub/TestShards.jl only; the lab-sotashimozono repos are a different account and were not measured. Worth measuring before anyone relies on cache: true, and it is a plausible part of why #24 found cache: false better on the persistent-depot runners.
    3. The git-ref transport is back on the table, and it deserves a fair re-read, because two of the three reasons I rejected it in [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80 were themselves wrong:
      • "it would newly require contents: write on jobs running arbitrary test code" — that is the status quo: the test job inherits the caller's grant, callers are required to grant it (:249-252), and :595 already injects secrets.GITHUB_TOKEN into the test process unconditionally.
      • "deleted refs leave unreachable objects" — true, but ci-timings is force-pushed every run by the same argument, so this is a size argument, not a principled one.
      • "a fork pull_request gets a read-only token" — this one stands, and is the only one that does. Under the visibility-derived mode it applies to public repos, which keep the artifact path anyway, so it constrains a narrow case rather than the design.

    What is actually left

    For shards → collect without artifact storage, after this measurement:

    • git refs — works, measured end to end locally; needs the fork-PR fallback
    • the Actions API, reading each shard's own job log from collect — actions: read, no storage, no writes, works on fork PRs; fragile, since it parses a log
    • cache — refuted above

    labels → shards via job output is unaffected by any of this and remains strictly better than today's artifact (903 B, fork-safe, and it makes the history-divergence failure at :290-293 structurally impossible).

    Reopening the transport question rather than the goal. The requirement — zero Actions storage, all five outputs — is unchanged.

  4. sotashimozono commented on Aug 20, 2026

    @sotashimozono
    MemberAuthor

    A transport that works, measured end to end: each shard emits its evidence into its own job log, and collect reads it back through the Actions API.

    run=32328961422 completed/success
    produce jobs: 96305777760 96305777898 96305777914
      recovered s3 from job 96305777760 (26 B)
      recovered s1 from job 96305777898 (26 B)
      recovered s2 from job 96305777914 (26 B)
    --- LOG transport recovered 3 of 3 ---
    OK-A: mid-run log read works; N->1 fan-in with NO storage and NO write token
    

    The go/no-go was whether a finished job's log is readable from a later job of the same, still in-progress run. It is.

    Why this is better than the artifact path it replaces, not merely equal

    artifacts (today) cache (#82 as proposed) job logs
    Actions storage metered — the whole problem none none
    works at all here yes no — refuted by measurement yes
    token needed runtime token runtime token actions: read only
    fork PR yes untested yes — actions: read is granted
    residue to clean up retention LRU none
    re-run isolation run-scoped (mixes attempts) key must carry run_attempt /attempts/{n}/jobs — exact

    It needs no write permission at all, which also disposes of the contents: write question that #80 got wrong in both directions.

    Four things that have to be right, each of which bit during the measurement

    1. gh api .../actions/jobs/{id}/logs does not follow the redirect. The endpoint answers 302 to a short-lived Azure blob URL; gh api exits 1 with no useful message. curl -sSL works. Measured:
      status=302 redirect=https://productionresultssa11.blob.core.windows.net/actions-results/.../job-logs.txt?...
      
    2. The runner echoes the step script into the log before running it, so a literal marker string appears twice — once as source, once as output. head -1 therefore reads back the unexpanded $(base64 -w0 ev.tsv) and the decode dies with base64: invalid input. Take the last match and exclude $ from the captured groups.
    3. Use /runs/{id}/attempts/{n}/jobs, not /runs/{id}/jobs, so a re-run reads its own attempt's jobs rather than merging attempts — the mirror of the flaw the cache design needed run_attempt in the key to avoid, and which today's artifacts (run-scoped, not attempt-scoped) actually have.
    4. Base64 the payload. A tab or a newline in a unit name would otherwise break the framing, and the marker has to be distinguishable from anything the suite itself prints.

    Payload sizes, and why #74 becomes a prerequisite for the coverage half

    One base64 line per shard:

    payload raw base64
    -out (shard-*.tsv, ran-*.tsv, timings, sections) 7–9 KB ~10–12 KB
    coverage counters, today ~104 KB ~140 KB
    coverage counters, after #74/#75's counter_index ~15 KB ~20 KB

    12 KB on one line is unremarkable; 140 KB is asking for trouble. So shipping #75's workflow side stops being an optimisation and becomes the thing that makes the coverage payload fit this channel — the two changes converge rather than compete.

    Also measured, and negative

    • actions/cache cannot carry this — see the previous comment. Save succeeds and the entry is listed as active; restore misses every key, including a default-branch entry written by this repo's own CI. Not a delay, not a scope rule.
    • The git-ref transport did not work inside Actions either: with contents: write granted and persist-credentials at its default, no ref reached the remote (git ls-remote origin 'refs/ci-run/*' was empty from the collector). Not debugged further, because the log transport needs no write token at all and is therefore strictly preferable. [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80's rejection of it stands, for better reasons than the ones it gave.

    Next

    Implement, in this order:

    1. Coverage (upload) — tolerate a missing -lcov. It is a name: download, which hard-fails on absence, while collect's pattern: download succeeds with zero results. One conditional; it turns a red job into a reported degradation.
    2. collect's completeness evidence — over the log channel, keeping the check itself untouched.
    3. Per-shard coverage artifacts ship 8 copies of the source tree — 85.7% of the payload is text collect already has checked out #74/Coverage counters can travel without the source they are printed against (#74) #75's wiring, then the coverage counters over the same channel.
    4. labels → shards via job output — unaffected by all of the above and still strictly better than its artifact.

    The probe branch and its caches have been deleted.

  5. sotashimozono commented on Aug 20, 2026

    @sotashimozono
    MemberAuthor

    Closing — keeping the current behaviour. The requirement was real and the transport question was answered; the answer is being left on the shelf rather than built.

    What is settled, and does not need re-measuring:

    question answer
    why artifacts stop Actions storage is accrued in GigabyteHours; CreateArtifact refuses before any bytes are sent, and deleting does not un-accrue the period
    actions/cache as transport does not work here — save succeeds and lists active, restore misses every key, including a default-branch entry and with a 45 s gap. Four runs, three controls
    job logs as transport works — recovered 3 of 3 in a live run; no storage, no residue, actions: read only, fork-PR safe
    git refs as transport pushes did not reach the remote from inside Actions; not debugged, since the log channel needs no write token
    job outputs fine for labels → shards (903 B); cannot serve shards → collect — matrix legs do not merge outputs, last leg wins

    And three implementation facts that cost a run each to learn:

    1. A new action and its first reference cannot land in one PR (uses: takes no expression; a relative reference in a reusable workflow resolves in the caller's checkout).
    2. gh api …/actions/jobs/{id}/logs answers 302 and gh does not follow it. curl -sSL does.
    3. The API names a reusable workflow's job with the caller's prefix — ci / test / shard s5. A name filter therefore depends on what the caller named its job, and one that matches nothing reads zero logs and reports a complete-looking empty run.

    The state this leaves: Coverage (upload) and the completeness gate stay red on every private consumer until the billing period rolls over. The gate failing is correct — it refuses rather than passing blind — and that is the behaviour being kept.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions