What kind
Cost grows where the model says it should not
Which benchmark
scripts/test-group.sh (the CI simulation groups)
The numbers
Per-test wall time on CI, median of five main runs on the seven-group
matrix, parsed from the job logs. Runs: 99135ec, 80b75eb, 89e73f9,
bb31393, c266c90, all ubuntu-latest at TEST_THREADS=4.
suite test median min max
sim_missile_live the_missile_check_draws_only_where_the_... 20.6m 12.0m 22.4m
sim_fall_live a_thing_falls_and_clips_the_way_the_eng... 19.8m 11.6m 22.3m
sim_hearing_live a_look_wakes_on_what_its_sector_heard 18.4m 13.4m 20.9m
sim_punch_live a_punch_lands_only_where_the_engine_lan... 17.6m 10.1m 18.0m
sim_player_damage_live a_claw_reaches_the_player_through_its_o... 16.5m 11.6m 18.0m
Each is one test. Its suite is that test plus a trivial one, except
sim_player_damage_live, which carries six.
All simulation work is 443.9 test-minutes. Over six groups at four threads
a perfect pack would be 18.5 minutes, but the floor is 20.6, set by
sim_missile_live's single test, because a test cannot be split across
threads.
The machine, and how quiet it was
GitHub's ubuntu-latest runners, one job per group, nothing else on the
runner. Per-test time across the five runs has a median coefficient of
variation of 17% and a maximum of 35%, which is the noise any measurement
here carries.
What you think is happening
A test that opens a session pays the tic statement's analysis once, about
three minutes on a standard runner, and then runs its tics. These five run
long because each one drives many tics through a single session rather than
because any one tic is slow.
The consequence for CI is structural rather than a matter of packing.
#547 packs the suites over six groups and every group lands between 19.8
and 20.6 minutes, which is where the longest single test puts the floor. No
assignment of suites to groups can go below it, and adding a seventh group
would not either: the total is already below the floor at six.
What would move it is splitting these tests so no single one is 20 minutes.
Two shapes worth weighing, neither obviously right:
- Split by tic range, so a suite's tics run as several tests against the
same seeded state. Each new test pays the analysis again, about three
minutes, so splitting one 20-minute test in two gives roughly 11.5 and
11.5 rather than 10 and 10, and the floor falls to about 12 minutes.
- Split by case rather than by tic, where a test covers several independent
seeded arms that could each be their own test. This costs the same
analysis per test and is only worth it where the arms are genuinely
independent.
Either way the gain is bounded by the analysis cost, so the question is how
many pieces are worth paying it for. Worth measuring before anyone commits
to a shape: what the analysis actually costs on a runner today, per
session, against what each of these tests spends on its tics.
This is not a regression. The tests have been this long since they were
written; the cost only became visible when the groups were packed on
measured times rather than on counts.
What kind
Cost grows where the model says it should not
Which benchmark
scripts/test-group.sh (the CI simulation groups)
The numbers
The machine, and how quiet it was
GitHub's ubuntu-latest runners, one job per group, nothing else on the
runner. Per-test time across the five runs has a median coefficient of
variation of 17% and a maximum of 35%, which is the noise any measurement
here carries.
What you think is happening
A test that opens a session pays the tic statement's analysis once, about
three minutes on a standard runner, and then runs its tics. These five run
long because each one drives many tics through a single session rather than
because any one tic is slow.
The consequence for CI is structural rather than a matter of packing.
#547 packs the suites over six groups and every group lands between 19.8
and 20.6 minutes, which is where the longest single test puts the floor. No
assignment of suites to groups can go below it, and adding a seventh group
would not either: the total is already below the floor at six.
What would move it is splitting these tests so no single one is 20 minutes.
Two shapes worth weighing, neither obviously right:
same seeded state. Each new test pays the analysis again, about three
minutes, so splitting one 20-minute test in two gives roughly 11.5 and
11.5 rather than 10 and 10, and the floor falls to about 12 minutes.
seeded arms that could each be their own test. This costs the same
analysis per test and is only worth it where the arms are genuinely
independent.
Either way the gain is bounded by the analysis cost, so the question is how
many pieces are worth paying it for. Worth measuring before anyone commits
to a shape: what the analysis actually costs on a runner today, per
session, against what each of these tests spends on its tics.
This is not a regression. The tests have been this long since they were
written; the cost only became visible when the groups were packed on
measured times rather than on counts.