Repository navigation
fix(test): make the iOS simulator smoke lane's red/green signal meaningful (#2491) - #3336
Conversation
… snapshot assertion (#3328) The lane asserted a regular snapshot must not carry snapshotQuality at all, but the product's contract is a disclosure, not an absence: the AX bridge publishes no verdict, and the XCTest runner the route falls back to always stamps its serving strategy. When the bridge probe circuit is open for the app generation (a legitimate cold/loaded-host state), the runner-served healthy 'tree' verdict made the lane red twice on main (runs 37145696803, 37459193672). The assertion now keys on the typed verdict: an absent (bridge) or healthy verdict is accepted, a degraded 'recovered' or 'sparse' verdict still fails, never on the warning string.
…tion was prevented (#2491) The replay steps in the iOS smoke job run with --retries 2 while the fixture-backed E2E step runs node --test with no retry layer, so one cold runner restart inside one 10s wait fails the whole job. Seven of the last ten main failures are this one class: a wait whose own error names wait_capture_stalled, wait_runner_restart_exhausted, or wait_readiness_exhausted and carries retriable: true — the product itself says 'retry' on those verdicts. The harness gains an opt-in per-step classifier; the iOS lane feeds it a predicate keyed on that typed conjunction only (wire retriable AND the wait taxonomy reason), never on error text. A re-attributed miss re-issues the step once; a readable miss (target absent, deadline exhausted after a readable capture, a wrong asserted value, a runner crash) stays a hard failure. Steps that already own a miss policy (allowFailure/expectFailure) are exempt so existing retry loops keep their semantics.
Size Report
Startup median (7 runs, lower is better):
|
There was a problem hiding this comment.
All reported issues were addressed across 7 files
Reply with feedback, questions, or to request a fix.
View guided diff | Turn on auto-fix | Re-trigger cubic
…ulary The acquisition assertion only rejected 'sparse' and unexempt 'recovered', so a verdict with any other state (renamed enum, diverged runner) fell through both branches and certified depth facts off a tree the lane never classified. The state set is now pinned to the kernel-declared SNAPSHOT_QUALITY_STATES, with a negative test for states outside it.
Coordinator pass actions (head now 98321a2)Cubic P2 (unrecognized
The red
Is #3328 addressed here or owed? Addressed on the assertion side; one completion item owed. #3328's required behavior is conditional: if the bridge-disabled fallback is legitimate for a regular snapshot, the assertion must assert the disclosed fallback through typed reasons instead of demanding absence — that is the decision this PR implements ( |
Re-run outcome (requested in the coordinator pass)Re-ran the lane at head
Checks on |
Coordinator pass 2 — disposition (head unchanged,
|
|
Re-run outcome (as promised): 37843633972 @ |
|
The code looks right at 98321a2, but I am calling this evidence-pending: I did not download the run artifacts, so I could not confirm that the recorded payloads replay through the new assertion, or whether the green re-run absorbed any re-issue. All 19 checks pass at 98321a2, and the Smoke Tests job, which runs the changed harness and the depth-frontier scenario, passed on re-run 37843633972. I did not check whether a failed Not blocking, take or leave: the re-issue predicate in live-harness.ts runs on every iOS step, including replay and batch steps, whose failures carry the nested wait's Is there a smaller design than this? I looked and found none: the re-issue layer is opt-in, sits on the shared harness, allows one extra attempt, and keys on typed reason plus retriable, while product-side retries inside wait would change the meaning of the reasons #2491 attributes. For the assertion to check the fallback disclosure as typed data, the snapshot route's circuit-disabled or bridge-fallback reason must first appear as a typed field in the snapshot response (#2972). |
Artifact confirmations and the unchecked gap (from run artifacts at
|
|
Thanks, this answers what I left open at 98321a2. The two recorded payloads pass the new assertion, the green re-run shows no re-issued step in its step history, and a failed |
|
Summary
Repairs two ruling defects that made the iOS simulator smoke lane's signal unreliable (attribution in #2491).
snapshotQualityassertion inassertSimulatorSnapshotTreeDepthFrontierdemanded an absence the product never promised. The product contract is a disclosure: the AX-bridge path publishes no verdict (the bridge has no quality model); the XCTest runner always stamps its strategy. A healthytreeverdict on the circuit-disabled fallback is legitimate acquisition, not a quality regression. Replaced withassertSimulatorSnapshotAcquisition, keyed on the typed verdict and reason code (rejectssparse; rejectsrecoveredunless pre-selecteddeferred/requested-backend), never on the warning text. Both recorded CI payloads (runs 37459193672, 37145696803) replay through the new assertion and pass.wait_runner_restart_exhausted,wait_capture_stalled,wait_readiness_exhausted): a shared harness seam (reattemptInfrastructureMiss) re-issues one failed live step, enabled by the iOS lane via a typed predicate onerror.retriable === trueAND a transport/observation wait reason. Hard failures stay red:wait_target_absent,wait_deadline_exceeded, wrong asserted values, runner crashes, non-wait failures. Parallel-run replays already had--retries 2; the E2E step had no layer. Field confirmation of the conjunction's narrowness on this PR's own lane: run 37843633972 hitwait_deadline_exceeded(24/25 readable polls) on the webview cold-start step and the harness declined to re-issue it, exactly the readable-miss case the set excludes; recorded in iOS smoke lane: two untracked failure signatures from the #2491 attribution pass (MAIN_THREAD_TIMEOUT on alert dismiss; replay timeout_cleanup_pending) #3337.Test/harness only — no production lines changed.
Validation
Head
ffd575ae7:pnpm check:quick— clean (lint + typecheck).pnpm test:integration:node— 147 tests, 135 pass, 0 fail, 12 skipped.node --teston the three touched test files — 28 pass, 0 fail.pnpm check:affected --run— all runnable checks passed.Unresolved risks: no live lane rerun performed (test-only change; CI re-runs the lane on this PR). Retried misses currently share one runner/session, so a runner-dead miss may re-fail — bounded at one re-issue; per-scenario fresh-runner teardown is the lane rewrite follow-up. Untracked new-signature failures observed during attribution (
MAIN_THREAD_TIMEOUTonalert dismissin 37660079219; gesture-replaytimeout_cleanup_pendingin 36115336398) need issue triage outside this scope.