Skip to content

Cut the artifact footprint that is actually there: retention on -out/-records, ship #75's workflow side, and stop leaking a write token to Pkg.test #81

Description

@sotashimozono

Replaces #80, whose premise was a measurement error (see its body). This is the small version, ordered by measured effect per line changed. Nothing here changes an input default, a permission, or the fork-PR path.

What the fleet actually stores

Live bytes only — filtering expired == false, because an expired artifact keeps its size_in_bytes metadata after GitHub deletes the bytes, and summing without that filter overcounts by every artifact ever created. Across the four largest consumers (86 of the account's 90.7 MB live):

family live files live bytes retention line
-out-* 7457 58.73 MB 7 days :649
-records 750 16.93 MB 30 days :844
-lcov 908 10.33 MB 7 days :794
-timings 98 0.11 MB 1 day :336
-coverage-* 0 0 MB 1 day :636
-prebuild 0 0 MB 1 day :429

Two families hold 84 % of it, and both are the ones whose retention is set for forensics rather than for the run.

1. Retention on -out-* and -records

Nothing in a run depends on either past its own run: collect consumes -out-* in the same run, timings consumes it in the same run, and -records is a 30-day merged archive nobody reads programmatically. At ~97 runs/day a 7-day window is ~680 live copies of data whose useful life is minutes.

Proposal: -out-* 7 → 2 days, -records 30 → 7 days. Both become inputs so a consumer who does want a forensic window can say so — defaults are the shortest value that keeps in-run consumption intact, which is what the other four families already do.

Cost: a shorter window for reading a failed shard's raw section timings after the fact. That is the whole cost, and it is a knob.

2. Ship the workflow side of #75

counter_index / restore_counters are merged, tested (test/core/test_coverage.jl) and not referenced anywhere in .github/ — the shard step at :611-624 still plain-cps the raw .cov. Measured on a real 48-file shard payload: 965,565 → ~15 KB, ~62×, round-trip byte-exact on 48/48 files, and restore_counters already accepts .cov.idx so old and new parts/ trees both restore.

This does not reduce live bytes (-coverage-* expires in a day and is already 0 MB live) — it reduces the bytes an in-flight run has to create, which is the thing a quota gate blocks on. ~5 lines.

3. Automate the prune

lab-sotashimozono/.github#32 already carries this. Demoted from "the fix" to what it is: a backstop. It cannot restore service inside a quota event, because usage is recalculated only every 6–12 hours — which is the one piece of #78's reasoning that survives, and the reason the recent outage lasted 18 hours after the underlying number was already fixed.

4. Move the labels → shards history onto a job output

needs.labels.outputs.timings already exists (:267, :482), added by #77 so the matrix cannot split into "some shards have history, some don't". timings.tsv is under 1 KB. Passing the content base64 through a job output instead of upload-artifact/download-artifact (:330-337 / :481-485):

  • removes the artifact whose failure gated the whole matrix during the outage
  • needs no new mechanism, no permission change, and works on fork PRs
  • hardens the invariant it carries: every shard provably reads the identical history, rather than each fetching ci-timings and hoping. That is the failure recorded at :290-293 — seven shards loaded 23 rows, the eighth loaded none, "two units ran twice, two ran nowhere, and only the completeness gate noticed" — made structurally impossible instead of merely detected.

Note the asymmetry that makes the other direction hard, and that #80 got wrong: GitHub does not merge job outputs across matrix legs. The last leg to finish overwrites, order unspecified. So shards → collect genuinely needs a transport, and the completeness evidence genuinely has to move. actions/cache is the candidate worth measuring (separate quota, LRU eviction rather than hard failure, one key per run_id/sid, restorable by exact key) — the #72 objection was about a multi-GB shared depot key, not about a 2 KB TSV under a unique one.

5. :595 hands a write-capable token to arbitrary test code unconditionally

Independent of everything above:

594          TESTSHARDS_CLAIM: ${{ inputs.steal && '1' || '' }}
595          TESTSHARDS_CLAIM_TOKEN: ${{ secrets.GITHUB_TOKEN }}

:594 is gated on inputs.steal. :595 is not. Since callers must grant contents: write (:249-252), that token is write-capable and sits in the environment of julia-actions/julia-runtest on every run of every consumer — readable by the suite and by every dependency it loads — whether claiming is enabled or not.

Fix: ${{ inputs.steal && secrets.GITHUB_TOKEN || '' }}, the same form :594 already uses. Note :582-584 documents in this file why the operand order matters (an empty string is falsy, so cond && '' || value sets it anyway).

Explicitly NOT proposed

  • Removing the artifact path. 90.7 MB live against a documented 500 MB pool; the storage is not structurally unusable, it was unmanaged. See [design] Artifact-free sharding on a free plan: move the completeness check to the shard instead of moving its data #80.
  • Per-shard Codecov uploads. Re-adds the lcov conversion :601-606 centralised on measured grounds (3–25 s per shard), multiplies upload count by 8 against another metered free-plan resource, and leaves a tokenless or fork-PR user with per-shard percentages — which do not average, and which this repo already recorded as "54.5 % for a suite that covered 94.8 %".
  • Weakening Every unit ran, exactly once. Failing closed when it cannot verify is correct for a gate. If it should survive an artifact failure, it needs an evidence path, not a lower bar.

Acceptance

  • live artifact bytes for a consumer measured before and after (1), filtering expired == false
  • a shard's .cov.idx restores byte-exact in collect, with a parts/ tree containing both formats
  • -timings no longer appears in the artifact list for a run, and all N shards report the same history row count
  • TESTSHARDS_CLAIM_TOKEN is empty in a steal: false run, asserted from the shard's own environment
  • a fork pull request still gets: tests, a completeness verdict, and a coverage number
  • shards: 1 and steal: true both still pass

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions