Skip to content

Audit the test suite for fixtures that assert shapes the real system never produces #2345

Description

@Brad-Edwards

Found in a sibling repository (Ground Control, issue #1683) and filed here
because the class is not specific to that codebase or that language — it is a
property of how boundary fixtures get written.

The failure class

A test fixture encodes a shape the real system never produces. The test passes,
the production code it guards agrees with the fixture, and the behaviour cannot
work. The suite stays green from the day the feature ships until something makes
first contact with the real thing.

This is not "a missing test." The test exists, it is well written, it names the
right behaviour, and it is wrong about the world.

The worked example

Ground Control's automated post-merge finalizer verified that a workflow run
belonged to a given pull request by reading the run's pull_requests
association from the GitHub Actions API. Its test fixture was:

const GOOD_RUN = {
  path: PHASE_E_WORKFLOW_PATH,
  event: "pull_request",
  pull_requests: [{ number: PR }],   // GitHub never returns this here
};

GitHub associates a workflow run with open pull requests only. The workflow
triggers on pull_request: closed, so on every real merge the API returns
pull_requests: []. The check was unsatisfiable on the only path the feature
existed to serve. It posted its report and then refused to close the issue,
every time, and had done so since it shipped.

Eleven tests covered that helper. All of them passed.

Why the usual signals miss it

  • Coverage says the line is covered. It is. Coverage measures execution, not
    whether the input was real.
  • Mutation testing inside the repository says the test is strong. Delete the
    production branch and the test goes red — because the fixture and the code
    share the same false belief. They are wrong together, so they confirm each
    other.
  • Code review misses it. The fixture looks plausible; the field exists in
    the API, it just is not populated in this state. Reviewing the diff does not
    tell you what the remote system returns.
  • Green CI is the whole problem. The signal that would normally reassure you
    is the signal that is lying.

What to look for

Every place a fixture stands in for something this repository does not control:

  • Third-party and platform API responses — including fields that exist but are
    conditionally populated, which is the exact trap above.
  • CLI or subprocess output: stdout shape, exit codes, stderr on failure, and
    what an empty or "no content" response actually looks like.
  • Database and driver return shapes, especially for empty results, deletes, and
    upserts.
  • Filesystem and OS behaviour: permissions, symlinks, path normalisation,
    what errors are actually raised.
  • Time, timezones, clock skew, and ordering.
  • Message, queue, and webhook payloads.
  • Anything mocked at a boundary where the mock's author inferred the contract
    from documentation, from a sibling call, or from what the code expected,
    rather than from a captured real response.

The sharpest question to ask of any fixture: where did this shape come from?
If the answer is "it is what the code expects," that fixture proves nothing.
If the answer is "it is what we observed the system return, in this state," it
is evidence.

Scope of this audit

  1. Inventory the boundaries this repository mocks or fakes in tests.
  2. For each, find the fixture and establish its provenance. Record which ones
    were captured from the real system and which were inferred.
  3. For the inferred ones, rank by blast radius: a wrong fixture on a security
    control, an authorization check, a data-integrity path, or a rarely exercised
    automated path is worse than one on a hot path that runs constantly and would
    fail loudly.
  4. Validate the ranked set against the real system. A single captured response
    is enough; the point is to replace a belief with an observation.
  5. Fix what turns out to be wrong, and fix the production code that agreed with
    it. Expect the fix to come in pairs.
  6. Leave the captured shapes in the repository so the next person inherits
    evidence rather than a new inference.

This is an audit, not a rewrite. Most fixtures will be fine. The deliverable is
knowing which ones were never checked, and having checked the ones that matter.

Acceptance criteria

  • A written inventory of the repository's test boundaries and, for each, whether
    its fixture shape was captured or inferred.
  • Every high-blast-radius inferred fixture validated against the real system,
    with the observed response recorded.
  • Every divergence found is fixed in both the fixture and the production code
    that relied on it, with a regression test that goes red on the real shape.
  • A short convention written down for how a new boundary fixture establishes its
    provenance, so this class does not re-enter through the next feature.
  • Findings that apply to sibling repositories are reported back, since this
    class is not repository-specific.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    tech-debttestingQuality assurance and testing tasks

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions