Skip to content

perf: five simulation suites are one twenty-minute test each #548

Description

@MarcusKainth

What kind

Cost grows where the model says it should not

Which benchmark

scripts/test-group.sh (the CI simulation groups)

The numbers

Per-test wall time on CI, median of five main runs on the seven-group
matrix, parsed from the job logs. Runs: 99135ec, 80b75eb, 89e73f9,
bb31393, c266c90, all ubuntu-latest at TEST_THREADS=4.

suite                     test                                        median    min    max
sim_missile_live          the_missile_check_draws_only_where_the_...   20.6m  12.0m  22.4m
sim_fall_live             a_thing_falls_and_clips_the_way_the_eng...   19.8m  11.6m  22.3m
sim_hearing_live          a_look_wakes_on_what_its_sector_heard        18.4m  13.4m  20.9m
sim_punch_live            a_punch_lands_only_where_the_engine_lan...   17.6m  10.1m  18.0m
sim_player_damage_live    a_claw_reaches_the_player_through_its_o...   16.5m  11.6m  18.0m

Each is one test. Its suite is that test plus a trivial one, except
sim_player_damage_live, which carries six.

All simulation work is 443.9 test-minutes. Over six groups at four threads
a perfect pack would be 18.5 minutes, but the floor is 20.6, set by
sim_missile_live's single test, because a test cannot be split across
threads.

The machine, and how quiet it was

GitHub's ubuntu-latest runners, one job per group, nothing else on the
runner. Per-test time across the five runs has a median coefficient of
variation of 17% and a maximum of 35%, which is the noise any measurement
here carries.

What you think is happening

A test that opens a session pays the tic statement's analysis once, about
three minutes on a standard runner, and then runs its tics. These five run
long because each one drives many tics through a single session rather than
because any one tic is slow.

The consequence for CI is structural rather than a matter of packing.
#547 packs the suites over six groups and every group lands between 19.8
and 20.6 minutes, which is where the longest single test puts the floor. No
assignment of suites to groups can go below it, and adding a seventh group
would not either: the total is already below the floor at six.

What would move it is splitting these tests so no single one is 20 minutes.
Two shapes worth weighing, neither obviously right:

  • Split by tic range, so a suite's tics run as several tests against the
    same seeded state. Each new test pays the analysis again, about three
    minutes, so splitting one 20-minute test in two gives roughly 11.5 and
    11.5 rather than 10 and 10, and the floor falls to about 12 minutes.
  • Split by case rather than by tic, where a test covers several independent
    seeded arms that could each be their own test. This costs the same
    analysis per test and is only worth it where the arms are genuinely
    independent.

Either way the gain is bounded by the analysis cost, so the question is how
many pieces are worth paying it for. Worth measuring before anyone commits
to a shape: what the analysis actually costs on a runner today, per
session, against what each of these tests spends on its tics.

This is not a regression. The tests have been this long since they were
written; the cost only became visible when the groups were packed on
measured times rather than on counts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceThroughput below expectation, or a regression

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions