Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 57 additions & 17 deletions .github/workflows/CI.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,17 +8,15 @@ concurrency:
group: ${{ github.workflow }}-${{ github.head_ref || github.ref }}
cancel-in-progress: true

# NOT the org's sharded-tests reusable, deliberately. That one shards a package's
# suite and expects TestShards in the test environment; this repository's library
# has one trivial test and would be forcing a dependency on everyone who presses
# "Use this template". What actually needs checking here is different anyway: that
# a project environment RESOLVES and that its scripts run — a green library beside
# a project that cannot instantiate would be the wrong kind of green.
# NOT the org's sharded-tests reusable, deliberately. That one shards a package's suite and expects
# TestShards in the test environment, which would force a dependency on everyone who presses "Use
# this template". What needs checking here is different: that every project environment resolves and
# that its smoke sweep runs. A green library beside a project that cannot instantiate would be the
# wrong kind of green.
#
# `setup.sh` is deliberately NOT run here. It rewrites names and re-issues UUIDs,
# which is a step the person starting a study should take knowingly rather than
# find already done for them. The cost is that its failure modes are not covered
# by CI, so a change to it has to be exercised by hand.
# `setup.sh` is deliberately NOT run here. It rewrites names and re-issues UUIDs, which is a step the
# person starting a study should take knowingly rather than find already done for them. The cost is
# that its failure modes are not covered by CI, so a change to it has to be exercised by hand.
jobs:
library:
runs-on: ubuntu-latest
Expand All @@ -30,20 +28,62 @@ jobs:
- uses: julia-actions/julia-buildpkg@v1
- uses: julia-actions/julia-runtest@v1

# Every directory under `projects/` that carries a Project.toml is its own environment, so each
# gets its own run rooted there. Discovered from the tree rather than listed here: a list and a
# tree disagree the first time someone copies a project and forgets to add it.
discover:
runs-on: ubuntu-latest
outputs:
projects: ${{ steps.find.outputs.projects }}
steps:
- uses: actions/checkout@v7
- id: find
run: |
set -euo pipefail
found="$(find projects -mindepth 2 -maxdepth 2 -name Project.toml -printf '%h\n' | sort)"
test -n "$found" || { echo "no projects/*/Project.toml found" >&2; exit 1; }
echo "$found" | sed 's/^/ /' >&2
echo "projects=$(printf '%s' "$found" | jq -R . | jq -sc .)" >> "$GITHUB_OUTPUT"

project:
needs: discover
strategy:
fail-fast: false
matrix:
project: ${{ fromJson(needs.discover.outputs.projects) }}
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: julia-actions/setup-julia@v3
with: { version: '1' }
- uses: julia-actions/cache@v3
- name: Instantiate the example project
working-directory: projects/ExampleSweep
- name: Instantiate
working-directory: ${{ matrix.project }}
run: julia --project=. -e 'using Pkg; Pkg.instantiate()'
- name: Its own tests
working-directory: projects/ExampleSweep
run: julia --project=. test/runtests.jl
- name: The reduction both faces call
working-directory: projects/ExampleSweep
run: julia --project=. scripts/collect.jl out
working-directory: ${{ matrix.project }}
run: |
set -euo pipefail
test -f test/runtests.jl || { echo "no test/runtests.jl" >&2; exit 1; }
julia --project=. test/runtests.jl
# Every project ships `configs/smoke.toml`: the size that proves the wiring before a queue
# does. Running it here is what makes this job mean "the sweep works", not "it resolved".
- name: Smoke sweep, then read it back
working-directory: ${{ matrix.project }}
run: |
set -euo pipefail
test -f configs/smoke.toml || { echo "no configs/smoke.toml" >&2; exit 1; }
julia --project=. scripts/compute.jl configs/smoke.toml
julia --project=. scripts/collect.jl configs/smoke.toml

# One stable context for the branch ruleset to require. A matrix job's own checks are named after
# the matrix entry, so requiring `project` directly would require a context that stops being
# reported the moment a project is added or renamed.
projects-passed:
needs: project
if: always()
runs-on: ubuntu-latest
steps:
- run: |
test "${{ needs.project.result }}" = "success" \
|| { echo "a project failed: ${{ needs.project.result }}" >&2; exit 1; }
6 changes: 1 addition & 5 deletions docs/make.jl
Original file line number Diff line number Diff line change
@@ -1,7 +1,3 @@
using Documenter, MyModule

makedocs(;
sitename="MyModule",
modules=[MyModule],
pages=["Home" => "index.md"],
)
makedocs(; sitename="MyModule", modules=[MyModule], pages=["Home" => "index.md"])
30 changes: 24 additions & 6 deletions projects/ExampleSweep/configs/debug.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,25 @@
# Small but not trivial — the size you submit once to see the queue behave.
[sweep]
L = [8, 12]
h = [0.5, 1.0, 1.5]
# One file, read by all three layers:
# ParamIO.load(this) -> ConfigSpec ([study], [datavault], [[paramsets]])
# DataVault.Vault(this) -> Vault ([study] for project_name/outdir,
# [datavault] path_keys for directory names)
# SweepRunner.run! -> over the DataKeys ParamIO expands from it
#
# A LIST value is a swept axis; a SCALAR is fixed. In DataKey.params the entries
# appear under their DOTTED names, e.g. "system.a".

[run]
out = "out"
[study]
project_name = "example"
total_samples = 1
outdir = "out"

[datavault]
# Which params name the on-disk directory and the canonical key identity.
path_keys = ["system.a", "numerics.dt"]

[[paramsets]]

[paramsets.system]
a = [0.5, 1.0]

[paramsets.numerics]
dt = [1.0e-2, 5.0e-3]
33 changes: 25 additions & 8 deletions projects/ExampleSweep/configs/production.toml
Original file line number Diff line number Diff line change
@@ -1,8 +1,25 @@
# The real sweep. Same file shape as `smoke.toml`; only the ranges differ, so a
# config that works small is the config that runs big.
[sweep]
L = [8, 12, 16, 20]
h = [0.2, 0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6, 1.8, 2.0]

[run]
out = "out"
# One file, read by all three layers:
# ParamIO.load(this) -> ConfigSpec ([study], [datavault], [[paramsets]])
# DataVault.Vault(this) -> Vault ([study] for project_name/outdir,
# [datavault] path_keys for directory names)
# SweepRunner.run! -> over the DataKeys ParamIO expands from it
#
# A LIST value is a swept axis; a SCALAR is fixed. In DataKey.params the entries
# appear under their DOTTED names, e.g. "system.a".

[study]
project_name = "example"
total_samples = 3
outdir = "out"

[datavault]
# Which params name the on-disk directory and the canonical key identity.
path_keys = ["system.a", "numerics.dt"]

[[paramsets]]

[paramsets.system]
a = [0.25, 0.5, 1.0, 2.0, 4.0]

[paramsets.numerics]
dt = [1.0e-2, 5.0e-3, 2.5e-3, 1.25e-3]
30 changes: 24 additions & 6 deletions projects/ExampleSweep/configs/smoke.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,25 @@
# Two points. Runs in seconds on a laptop; proves the wiring before a queue does.
[sweep]
L = [8]
h = [0.5, 1.0]
# One file, read by all three layers:
# ParamIO.load(this) -> ConfigSpec ([study], [datavault], [[paramsets]])
# DataVault.Vault(this) -> Vault ([study] for project_name/outdir,
# [datavault] path_keys for directory names)
# SweepRunner.run! -> over the DataKeys ParamIO expands from it
#
# A LIST value is a swept axis; a SCALAR is fixed. In DataKey.params the entries
# appear under their DOTTED names, e.g. "system.a".

[run]
out = "out"
[study]
project_name = "example"
total_samples = 1
outdir = "out"

[datavault]
# Which params name the on-disk directory and the canonical key identity.
path_keys = ["system.a", "numerics.dt"]

[[paramsets]]

[paramsets.system]
a = [1.0]

[paramsets.numerics]
dt = [1.0e-2]
21 changes: 15 additions & 6 deletions projects/ExampleSweep/report/report.jl
Original file line number Diff line number Diff line change
@@ -1,8 +1,17 @@
# Draw from the finished vault. Run with `--project=report`, never the compute env.
# Draw from the finished vault. Run with `--project=report`, never the compute env —
# that is the whole reason the two environments are separate.
#
# julia --project=report report/report.jl out
using DataVault, ExampleSweep
# julia --project=report report/report.jl configs/smoke.toml

vault = DataVault.Vault(get(ARGS, 1, "out"))
s = ExampleSweep.summarise(vault) # the same reduction scripts/collect.jl uses
@info "reporting over" s...
using DataVault: DataVault
using ExampleSweep: ExampleSweep
using ParamIO: ParamIO

const CONFIG = get(ARGS, 1, joinpath(@__DIR__, "..", "configs", "smoke.toml"))
const OUTDIR = get(ENV, "DATAVAULT_OUTDIR", joinpath(@__DIR__, "..", "out"))

vault = DataVault.Vault(CONFIG; run="phase1", outdir=OUTDIR)
rows = ExampleSweep.summarise(vault) # the same reduction scripts/collect.jl uses

@info "reporting over" n = length(rows)
# Draw here — Pinax and a plotting backend are dependencies of THIS environment.
27 changes: 20 additions & 7 deletions projects/ExampleSweep/scripts/collect.jl
Original file line number Diff line number Diff line change
@@ -1,9 +1,22 @@
# Read the finished vault and report it as text. No plotting dependency, so this
# can run on the machine that did the compute.
# The reader side. `compute.jl` never called save! — the runtime persisted every
# work_fn return for us. No plotting dependency, so this runs where the compute did.
#
# julia --project=. scripts/collect.jl out
using DataVault, ExampleSweep
# julia --project=. scripts/collect.jl configs/smoke.toml

vault = DataVault.Vault(get(ARGS, 1, "out"))
s = ExampleSweep.summarise(vault) # the same reduction report/report.jl uses
@info "collected" s...
using DataVault: DataVault
using ExampleSweep: ExampleSweep
using ParamIO: ParamIO
using Printf

const CONFIG = get(ARGS, 1, joinpath(@__DIR__, "..", "configs", "smoke.toml"))
const OUTDIR = get(ENV, "DATAVAULT_OUTDIR", joinpath(@__DIR__, "..", "out"))

vault = DataVault.Vault(CONFIG; run="phase1", outdir=OUTDIR)
rows = ExampleSweep.summarise(vault) # the same reduction report/ uses

@printf("\n %-8s %-10s %s\n", "a", "dt", "rel_error")
println(" ─────────────────────────────────────")
for r in rows
@printf(" %-8.4g %-10.4g %.3e\n", r.a, r.dt, r.rel_error)
end
@printf("\n %d points\n\n", length(rows))
66 changes: 46 additions & 20 deletions projects/ExampleSweep/scripts/compute.jl
Original file line number Diff line number Diff line change
@@ -1,20 +1,46 @@
# One entry point for every size of run. `julia --project=. scripts/compute.jl configs/smoke.toml`
#
# What SweepRunner adds over a `for` loop: a point already finished is skipped
# after one manifest read, two processes pointed at the same vault never compute
# the same point twice, and a killed run is continued rather than restarted.

using DataVault
using MyModule
using ParamIO
using SweepRunner

config = get(ARGS, 1, "configs/smoke.toml")
spec = ParamIO.load(config)
vault = DataVault.Vault(spec["run"]["out"])
keys = ParamIO.enumerate_keys(spec["sweep"])

work_fn(key) = MyModule.solve(; ParamIO.params(key)...)

result = SweepRunner.run!(work_fn, vault, keys; load = MyModule)
SweepRunner.launchable(result) || exit(1)
#==============================================================================
compute.jl — the three-layer driver. This is the file to copy, not to invent.

ParamIO : config TOML -> Vector{DataKey} (what to compute)
DataVault : (study, run) -> file storage (where it goes)
SweepRunner: run!(work_fn) -> parallel runtime (do it, lock-safe, resumable)

The work itself is `MyModule.work_fn`, in a PACKAGE rather than in this script.
That is what lets `run!(…; load=MyModule)` hand it to the workers — no
`@everywhere`, and nothing to forget broadcasting.

julia --project=. scripts/compute.jl configs/smoke.toml

Run it again and it exits in milliseconds: the manifest records what is done.
==============================================================================#

# `using X: X` brings in the MODULE and none of its exports, so every call below has to name where
# it comes from. That is a discipline the seam needs rather than a style choice: ClassicalMonteCarlo
# — a plausible work package for this slot — also exports `run!`, and a bare `using` of both would
# make `run!` ambiguous at the one call that matters. Qualifying keeps `SweepRunner.run!` the sweep's
# `run!` no matter what the work package is called or what it exports.
using DataVault: DataVault
using MyModule: MyModule
using ParamIO: ParamIO
using SweepRunner: SweepRunner

const CONFIG = get(ARGS, 1, joinpath(@__DIR__, "..", "configs", "smoke.toml"))
# outdir precedence, resolved by DataVault: kwarg > ENV > the config's [study].
const OUTDIR = get(ENV, "DATAVAULT_OUTDIR", joinpath(@__DIR__, "..", "out"))

spec = ParamIO.load(CONFIG)
keys = ParamIO.expand(spec)
vault = DataVault.Vault(CONFIG; run="phase1", outdir=OUTDIR)

SweepRunner.init_workers!(; mode=:auto)

# work_fn RETURNS a Dict; the runtime saves it and writes the .done marker.
# batch/run.sh traps the wall-clock signal and touches this file, at which point
# run! stops dispatching new keys and returns cleanly instead of being killed.
opts = SweepRunner.RunOpts(; stop_flag=get(ENV, "PM_STOP_FLAG", nothing))

result = SweepRunner.run!(MyModule.work_fn, vault, keys; opts=opts, load=MyModule)
@info "phase1 complete" result

ledger = DataVault.build_ledger(vault) # one row per completed key
@info "ledger written" ledger
24 changes: 13 additions & 11 deletions projects/ExampleSweep/src/ExampleSweep.jl
Original file line number Diff line number Diff line change
@@ -1,24 +1,26 @@
module ExampleSweep

# Project-local code: how THIS study reduces its own results.
#
# It lives here, and not in `scripts/` or in `report/`, because both of those
# call it. A summary printed on the cluster and a figure drawn afterwards then
# cannot report different numbers — there is only one function to be wrong.
# This study's own reduction. It lives here, and not in `scripts/` or `report/`,
# because both call it — a summary printed on the cluster and a figure drawn
# afterwards then cannot report different numbers.

using DataVault
using DataVault: DataVault

export summarise

"""
summarise(vault) -> NamedTuple
summarise(vault) -> Vector{NamedTuple}

Reduce a finished vault to whatever this study reports. Replace the body; keep
the shape, so `scripts/collect.jl` and `report/report.jl` stay in agreement.
Read every finished point back and reduce it. Replace the body; keep the shape,
so `scripts/collect.jl` and `report/report.jl` stay in agreement.
"""
function summarise(vault)
ks = DataVault.keys(vault)
return (; n_points=length(ks))
rows = NamedTuple[]
for key in DataVault.keys(vault; status=:done)
d = DataVault.load(vault, key)
push!(rows, (; a=d["a"], dt=d["dt"], rel_error=d["rel_error"]))
end
return sort(rows; by=r -> (r.a, r.dt))
end

end # module ExampleSweep
Loading
Loading