Found in a sibling repository (Ground Control, issue #1683) and filed here
because the class is not specific to that codebase or that language — it is a
property of how boundary fixtures get written.
The failure class
A test fixture encodes a shape the real system never produces. The test passes,
the production code it guards agrees with the fixture, and the behaviour cannot
work. The suite stays green from the day the feature ships until something makes
first contact with the real thing.
This is not "a missing test." The test exists, it is well written, it names the
right behaviour, and it is wrong about the world.
The worked example
Ground Control's automated post-merge finalizer verified that a workflow run
belonged to a given pull request by reading the run's pull_requests
association from the GitHub Actions API. Its test fixture was:
const GOOD_RUN = {
path: PHASE_E_WORKFLOW_PATH,
event: "pull_request",
pull_requests: [{ number: PR }], // GitHub never returns this here
};
GitHub associates a workflow run with open pull requests only. The workflow
triggers on pull_request: closed, so on every real merge the API returns
pull_requests: []. The check was unsatisfiable on the only path the feature
existed to serve. It posted its report and then refused to close the issue,
every time, and had done so since it shipped.
Eleven tests covered that helper. All of them passed.
Why the usual signals miss it
- Coverage says the line is covered. It is. Coverage measures execution, not
whether the input was real.
- Mutation testing inside the repository says the test is strong. Delete the
production branch and the test goes red — because the fixture and the code
share the same false belief. They are wrong together, so they confirm each
other.
- Code review misses it. The fixture looks plausible; the field exists in
the API, it just is not populated in this state. Reviewing the diff does not
tell you what the remote system returns.
- Green CI is the whole problem. The signal that would normally reassure you
is the signal that is lying.
What to look for
Every place a fixture stands in for something this repository does not control:
- Third-party and platform API responses — including fields that exist but are
conditionally populated, which is the exact trap above.
- CLI or subprocess output: stdout shape, exit codes, stderr on failure, and
what an empty or "no content" response actually looks like.
- Database and driver return shapes, especially for empty results, deletes, and
upserts.
- Filesystem and OS behaviour: permissions, symlinks, path normalisation,
what errors are actually raised.
- Time, timezones, clock skew, and ordering.
- Message, queue, and webhook payloads.
- Anything mocked at a boundary where the mock's author inferred the contract
from documentation, from a sibling call, or from what the code expected,
rather than from a captured real response.
The sharpest question to ask of any fixture: where did this shape come from?
If the answer is "it is what the code expects," that fixture proves nothing.
If the answer is "it is what we observed the system return, in this state," it
is evidence.
Scope of this audit
- Inventory the boundaries this repository mocks or fakes in tests.
- For each, find the fixture and establish its provenance. Record which ones
were captured from the real system and which were inferred.
- For the inferred ones, rank by blast radius: a wrong fixture on a security
control, an authorization check, a data-integrity path, or a rarely exercised
automated path is worse than one on a hot path that runs constantly and would
fail loudly.
- Validate the ranked set against the real system. A single captured response
is enough; the point is to replace a belief with an observation.
- Fix what turns out to be wrong, and fix the production code that agreed with
it. Expect the fix to come in pairs.
- Leave the captured shapes in the repository so the next person inherits
evidence rather than a new inference.
This is an audit, not a rewrite. Most fixtures will be fine. The deliverable is
knowing which ones were never checked, and having checked the ones that matter.
Acceptance criteria
- A written inventory of the repository's test boundaries and, for each, whether
its fixture shape was captured or inferred.
- Every high-blast-radius inferred fixture validated against the real system,
with the observed response recorded.
- Every divergence found is fixed in both the fixture and the production code
that relied on it, with a regression test that goes red on the real shape.
- A short convention written down for how a new boundary fixture establishes its
provenance, so this class does not re-enter through the next feature.
- Findings that apply to sibling repositories are reported back, since this
class is not repository-specific.
Found in a sibling repository (Ground Control, issue #1683) and filed here
because the class is not specific to that codebase or that language — it is a
property of how boundary fixtures get written.
The failure class
A test fixture encodes a shape the real system never produces. The test passes,
the production code it guards agrees with the fixture, and the behaviour cannot
work. The suite stays green from the day the feature ships until something makes
first contact with the real thing.
This is not "a missing test." The test exists, it is well written, it names the
right behaviour, and it is wrong about the world.
The worked example
Ground Control's automated post-merge finalizer verified that a workflow run
belonged to a given pull request by reading the run's
pull_requestsassociation from the GitHub Actions API. Its test fixture was:
GitHub associates a workflow run with open pull requests only. The workflow
triggers on
pull_request: closed, so on every real merge the API returnspull_requests: []. The check was unsatisfiable on the only path the featureexisted to serve. It posted its report and then refused to close the issue,
every time, and had done so since it shipped.
Eleven tests covered that helper. All of them passed.
Why the usual signals miss it
whether the input was real.
production branch and the test goes red — because the fixture and the code
share the same false belief. They are wrong together, so they confirm each
other.
the API, it just is not populated in this state. Reviewing the diff does not
tell you what the remote system returns.
is the signal that is lying.
What to look for
Every place a fixture stands in for something this repository does not control:
conditionally populated, which is the exact trap above.
what an empty or "no content" response actually looks like.
upserts.
what errors are actually raised.
from documentation, from a sibling call, or from what the code expected,
rather than from a captured real response.
The sharpest question to ask of any fixture: where did this shape come from?
If the answer is "it is what the code expects," that fixture proves nothing.
If the answer is "it is what we observed the system return, in this state," it
is evidence.
Scope of this audit
were captured from the real system and which were inferred.
control, an authorization check, a data-integrity path, or a rarely exercised
automated path is worse than one on a hot path that runs constantly and would
fail loudly.
is enough; the point is to replace a belief with an observation.
it. Expect the fix to come in pairs.
evidence rather than a new inference.
This is an audit, not a rewrite. Most fixtures will be fine. The deliverable is
knowing which ones were never checked, and having checked the ones that matter.
Acceptance criteria
its fixture shape was captured or inferred.
with the observed response recorded.
that relied on it, with a regression test that goes red on the real shape.
provenance, so this class does not re-enter through the next feature.
class is not repository-specific.