Skip to content

chore: sync with tester-army/e2e main - #3

Merged
amankansal-lt merged 51 commits into
LambdaTest:mainfrom
amankansal-lt:chore/sync-upstream-main
Oct 5, 2026
Merged

amankansal-lt merged 51 commits into
LambdaTest:mainfrom
amankansal-lt:chore/sync-upstream-main

Conversation

@amankansal-lt

Copy link
Copy Markdown
Collaborator

Summary

Brings LambdaTest/e2e main up to date with tester-army/e2e main (49 commits) so that tester-army#772, whose head is this fork's main, stops conflicting.

Please merge this with a merge commit, not squash or rebase, so main fast-forwards and tester-army#772 picks up upstream's history.

  • 8773058b merges tester-army main. Conflicts were in README.md, docs/package.json and package.json. In each, upstream's version is kept and only the testmu additions are re-added: the README row, the docs devDependency, and the build/test scripts.
  • 4a6def63 brings @e2e-dev/testmu in line with upstream:
    • agent-device devDependency 0.21.18 → 0.21.20, matching @e2e-dev/mobile's pin;
    • peer >0.21.20 <1;
    • Node engines ^22.22.3 || >=24.8.0;
    • the regenerated lockfile. The merged lockfile still pointed testmu at 0.21.18, which upstream had removed, so CI would have failed on a clean cache.

No testmu source changes were needed: upstream didn't touch the DeviceProvider/record() contract or the docs structure testmu uses.

Validation

Run on Node 22.22.3:

  • lint, check:dead-code, typecheck, docs:check-errors, check:peer-ranges, check:install-scripts, docs:check and test all pass. testmu: 73 tests; mobile: 232; e2e: 3160.
  • git merge-tree against tester-army main is clean.

🤖 Generated with Claude Code

okwasniewski and others added 30 commits October 2, 2026 18:17
…share a comment (tester-army#720)

* fix(github): long marker fields keep a digest so distinct keys never share a comment

* docs(changeset): note one-time duplicate for long marker fields
…ent input (tester-army#727)

* fix(e2e): redact plain-string secret values in titles, labels, and agent input

* fix(e2e): redact explore goal, keep unique() templates aligned with redacted params

* test(e2e): use literal escaped marker in junit assertion

* fix(e2e): refuse param keys that redact alike, bound explore goal after redaction

* fix(e2e): redact titles before NFC, drop duplicate-title heuristic, clarify docs

* fix(e2e): redact agent context per step, repair text before cut, re-check title limit

* fix(e2e): keep __proto__ param key when redacting, scope params wording to act
…wright does (tester-army#770)

* fix(web): list an inline svg as an image, named by its title, as Playwright does

Fixes tester-army#749

* test(web): count only compared crosscheck nodes, fix svg fixture comment
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* docs(mobile): document physical devices

Closes tester-army#766

* docs(mobile): phone requirements in setup and README

* docs(mobile): move physical devices to its own page
…roperty primitives, toHaveURL ignoreCase (tester-army#726)

Rebuilt on top of tester-army#579, which already landed matcher option flags, ignoreCase, and unknown-option refusal.
…y#713)

* fix(report)!: count interrupted tests apart from failures

* fix(report): tell a failed test from its last failed attempt, whatever the cut retry left
…r-army#777)

A route segment of letters, a dash, and a digit tail of two or more
(`PROJ-016`, `INV-2041`) now abstracts to `:id` like the other minted
shapes. These human-readable record ids stayed literals, so a step
recorded on one record missed the start-path precondition on the next
record and ran live every time.

A single trailing digit (`page-2`) stays a route word.

Fixes tester-army#769
…r-army#781)

* docs(models): add an AI SDK provider picker to the Models page

Readers did not see that e2e takes a model from any AI SDK provider. The
Models page now has a provider dropdown with 51 providers from
ai-sdk.dev, and three setup steps that follow the selection: install,
environment, and config.

The steps are four instances of one snippet component. Mintlify
evaluates each snippet export alone, so the instances share the
selection through a store on globalThis and React.useSyncExternalStore.
Code renders with Mintlify's CodeBlock for syntax highlighting.

* docs(models): an unlisted provider's model needs tool calls and images
* fix(runner): preserve failures when tests skip

* refactor(runner): one skipped-after-failure helper, evidence from failing attempt

- failureBeforeSkip returns failing attempt's error, evidence, artifacts; list reporter prints those instead of skipped attempt's
- someSkippedAfterFailure indexes serial groups once, shared by exit code and run status
- collapse skip branch in execute, place failOnSkippedFailure next to retries with JSDoc

---------

Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…aunches (tester-army#711)

* fix(mobile): never pin an install's file path as the app app.open() launches

agent-device 0.21.18 can return an Android install with no package (its binary manifest parser misses aapt2 output, and the aapt fallback needs ANDROID_HOME), echoing the APK path as app. The engine pinned that path, so app.open() failed later with 'Android runtime hints require an installed package name'. installApp now opens by the reported id or the app passed in, and fails with ENGINE_FAILURE when there is neither. app.open() keeps the engine's own UNSUPPORTED_CAPABILITY reason instead of a generic rewrite.

* test(mobile): pin the reported package, not the echoed path; clarify installApp failure docs
tester-army#721)

* fix(config): refuse secrets and credentials sharing an override env var

* docs(environment): name both credential override vars in collision rule
…ion (tester-army#722)

* fix(security): allow only http(s) navigation, refuse wrapped schemes on device links

view-source:file:// passed the file:/data:/javascript: denylist and loaded local files via app.open, browser.goto, and the navigate verb. resolveNavigationUrl now admits http: and https: only; mobile links also refuse view-source:, blob:, filesystem:.

* fix(security): admit exact about:blank, sync rule text in AGENTS.md and SECURITY.md

* fix(web): refuse about:blank cookie URL, qualify scheme rule text
…bedded names (tester-army#728)

* fix(web): read visibility as rendered, null box, checkable values, embedded names

* fix(web): narrow visible queries by the reader's hidden state

* fix(web): drop direct text a content-visibility hidden element skips

* docs(changeset): note the display contents divergence from Playwright
…ester-army#779)

* fix(oauth): reach Copilot models served only over the Responses API

copilot() reads the plan's model listing on the first call and picks
chat completions or the Responses API for each model, falling back to
chat whenever the listing cannot be read so nothing that worked changes.
sendCopilotRequest reads the initiator and vision headers from a
Responses body (input[]/input_image) as well as a chat one, and the model
listing marks an uncallable model 'no chat or responses'.

* fix(changeset): schedule a minor release for the Responses endpoint support

The changeset asked for a patch while the change adds a call path a model could not
reach before; a minor bump is what the release intent states.

Addresses the cubic review comment on .changeset/copilot-responses-endpoints.md.

* fix(oauth): cancel and never cache a Copilot endpoint lookup that was aborted

The first call read the plan's model listing with no signal, and the delegate cached
whatever that produced. A listing that stalled therefore left `pending` unresolved and
every retry waited on the same promise instead of reaching the model. The signal now
reaches the lookup, an aborted lookup rejects instead of being read as a missing
endpoint, and a choice that could not be made clears `pending` so the next attempt
asks again. A test pins the aborted lookup and the missing-endpoint read as distinct.

Addresses the cubic review comment on packages/e2e/src/oauth/copilot.ts.

* docs(skill): separate Copilot models served over Responses from uncallable ones

The setup reference read as if every non-chat model is reached over the Responses API.
Only an enabled model the plan serves solely over `/responses` is; a model served over
an API `copilot()` does not speak, or one the plan has not enabled, still cannot be
called, and the listing marks those.

Addresses the cubic review comment on skills/e2e/references/setup.md.

* fix(oauth): pick Copilot endpoint from supported_endpoints alone, retry unreadable listing

- drop policy gate: grok-4.5, gpt-5.3-codex serve /responses with no policy
- unreadable listing: call over chat, not cached, asked again next call
- listing has own 10s timeout; one caller abort no longer fails others
- revoked login fails on listing with sign-in hint, no second 401
- missing @ai-sdk/openai names the package to install

* fix(init): install @ai-sdk/openai for the Copilot gateway

* docs(copilot): unreadable listing retry, init installs both peers

---------

Co-authored-by: Hermes Agent <agent@localhost>
Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com>
Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ywright options (tester-army#723)

* fix(web): continue skips earlier routes, decisions take or reject Playwright options

continue went through Playwright fallback, so an earlier route answered. It now continues to the network, merging site headers itself; fallback() keeps chaining. continue takes url/method/headers/postData (url via resolveUrl), fulfill takes path/contentType; unknown options are INVALID_ARGUMENT. waitForResponse timeout bounds the match only; setCookies passes the resolved url.

* fix(web): route header overrides follow the web({ headers }) grammar

* refactor(web): route options reuse rejectUnknownOptions

* test(web): route url fake follows the http(s) navigation allowlist

* fix(web): fulfill keeps a content-type header over the json or path type
* feat(agent): identify e2e on every model request

* docs(changeset): note attribution headers replace provider-level ones

* test(agent): check headers on a second act-loop turn
…ts toward negation window (tester-army#725)

* fix(expect): unawaited poll fails its own phase, slow first read counts toward negation window

- expect.poll is owned by the phase that started it (body, hook, fixture teardown, suite hook); left running it is cancelled and fails that phase with STEP_NOT_AWAITED instead of a later test
- pollCondition starts the negation window when the first read that saw it was issued and decides at the deadline without a read past it

* fix(expect): own polls by async lineage so a timed-out body never charges the next test

* docs(expect): scope STEP_NOT_AWAITED wording to where it surfaces; test reads past body
…tester-army#729)

* fix(e2e): redact every observed node field and collapsed cut secrets

Observation text, executor trees, and cache anchors leaked a secret held in a test id, frame path, selector, or attribute, and the leading part of a secret with whitespace an engine collapsed before cutting. Nodes now pass one exhaustive redactNode before any reader, and redactCut and redactFragments compare in collapsed-whitespace form.

* perf(e2e): skip second redaction pass for already-collapsed fields

* perf(e2e): build collapsed reading key in typed arrays, fast path when nothing collapses

* refactor(e2e): descriptors, anchors, and relocation read only redacted nodes

describeTarget, containerKey, describeNodes, relocateDescriptor, describePosition, describeAnchors, and anchorsPresent take RedactedNode and drop their own weaker redact pass; ReplayHost.redact goes with it.

* fix(e2e): re-redact baseline nodes with the current ledger before deriving anchors

* test(e2e): count typed-array memory in the large-text redaction guard

* fix(e2e): re-redact an action's nodes with the current ledger when describing it

* fix(e2e): re-redact an action's container key with the current ledger
…ied debts, kept evidence (tester-army#724)

* fix(run): keep --last-failed reruns honest: failure-limit skips, carried debts, kept evidence

--last-failed now reruns tests --max-failures skipped. A test or failed
suite hook another filter leaves out stays owed in the report's new
run.carried until a rerun runs it; the github fold shows carried rows and
errors and stays red while anything is carried. A rerun prunes artifacts/
to what the previous report names and writes its attempts under
artifacts/rerun-<n>/ instead of wiping the tree and reusing attempt paths.

* fix(run): carry a setup's hook debt when no selected test needs it

The report has no row for an unneeded setup, so carryForward now asks the
collection whether a test still exists instead of the report's rows.

* fix(run): reject malformed artifacts in --last-failed reports, tighten rerun docs

Distinct ids in the carried fixture row.

* fix(run): keep carrying tests of a file a narrowed rerun could not collect

* fix(github): a carried interrupted test reads interrupted, not failed, in the folded page
…en (tester-army#732)

* fix(cache): stop replays self-finalizing when the effect did not happen

Key entries by agent name and a digest of the redacted agent context.
Routes carry origin and query, minted values as ids; the end route must
match exactly. Anchors record what vanished, control states and values,
announcements first; a replay needs evidence the delta happened since the
first screen on its end route, and an unseen alert fails it. Unnamed twins
need a named row. Steps with no observable delta are not recorded.

Policy bumped to conservative/6: every entry re-records.

* fix(cache): count call occurrences per agent

* fix(cache): hold the end route through the end wait, tighten review nits

* fix(cache): fold the agent context for occurrence counts like the key does

* refactor(cache): key on the agent context dispatch already redacted

* test(e2e): earlier capture redacted again keeps its tree shape
…ne every attempt runs in (tester-army#786)

* feat(web): locale and timezoneId options set the language and time zone every attempt runs in

* docs(web): timezoneId validation is best effort, the browser has the final word
…rmy#793)

device.openLink built its INVALID_ARGUMENT message from the raw input, so a
string that fails to parse as an absolute URL still echoed whatever it
carried. A malformed magic link is exactly the case where a one-time token
sits in that string, and the message reaches logs, the report, and CI
output.

Name no part of the input, the same rule assertAppId already follows in this
file: nothing about a string that failed to parse says which part is safe to
repeat, and a token can sit in the query, the path, or the userinfo. linkLabel
stays for the step label and the report, where the link did parse.

Covered by a unit test over malformed links whose token sits in the query,
the path, and the userinfo.
…tester-army#788)

* fix(web): a check whose control the click replaces passes and replays

check and uncheck click once and read the state back on the same element.
A control gone by that read (replaced, or navigated away) took the click,
so the action is done instead of NODE_STALE, which the agent turned into
LOCATOR_NOT_FOUND and left out of its recording.

Fixes ENG-814

* docs(changeset): name the checked radio uncheck refuses
…ter-army#794)

* feat(oauth): sign in with OpenCode Console for Zen and Go models

* fix(oauth): keep the Console refresh token, isolate tests from OPENCODE_API_KEY

A refresh grant without a refresh_token now keeps the previous one, as
the ChatGPT and SpaceXAI providers do. The vendor test helper blanks
OPENCODE_API_KEY so a key in the shell no longer replaces the stored
login under test. Docs and changeset now say the models listing covers
Zen and Go only, and that MISCONFIGURED needs a readable workspace config.
…6 cache key (tester-army#796)

tester-army#732 bumped REPLAY_POLICY_VERSION to conservative/6 and keyed entries by
agent, orphaning every committed entry. Re-recorded tests-agent live and
removed the 29 orphaned entries.
…ster-army#798)

* feat(web)!: pin playwright-core and add an install command

@e2e-dev/web depends on playwright-core 1.63.0 instead of peering on
playwright. New e2e-web bin installs browsers for the pinned version:
npx @e2e-dev/web install chromium --with-deps. e2e init no longer adds
playwright.

* test(web): run the e2e-web bin for real, and in CI

* test(web): spawn the built e2e-web bin so Node 22.12 runs it
thymikee and others added 21 commits October 3, 2026 22:23
…ipts }) or browser.addInitScript (tester-army#795)

* feat(web): init scripts run before the page's own, configured with web({ initScripts }) or added per test with browser.addInitScript

* refactor(web): one owner for configured init scripts, read beside the browser launch

* feat(web): export WebInitScript

* fix(web): settle the browser launch before a failed init-script read throws; docs examples compile

* fix(web): refuse sparse initScripts, keep an own __proto__ key in the init-script argument

* docs(web): seed Math.random in the init-script examples

* docs(web): trim init-script docs to what users need

* docs(skill): addInitScript takes arg with a function only
…anchor (tester-army#799)

* fix(cache): record a vanished node only when no end node matches its anchor

describeDelta keyed the gone side exactly while deltaHolds matches on the recorded fields alone, so a radio label {text: Express} the pick replaced with a status reading Express was recorded as gone and failed its own check on every replay.

* test(web-benchmark): record the radio-replaced-by-summary step
…ecording (tester-army#800)

* fix(cache): --strict-cache fails a step whose key changed under its recording

A cache key change (REPLAY_POLICY_VERSION, a new key field, an engine minor, the agent's context) turned every committed entry into a no-entry miss, which strict ran live. Strict now lists the file store once per run and fails a no-entry step with REPLAY_STALE when an entry recorded for the same step sits under another key. Entries now record the step's params digest, occurrence, and agent so a repeat or another agent is not mistaken for it.

* fix(cache): never take an entry the attempt claimed for a rekeyed step

An entry from before the occurrence fields matches by test, target, and instruction; an earlier occurrence's entry is now excluded, so an unrecorded repeat runs live. The message no longer says to delete the old entry, which may still serve another app identity.

* fix(cache): match a rekeyed recording only by its whole step

An entry without the params, occurrence, and agent could be another call of the same instruction, so strict now ignores it instead of guessing. A confirmed whole replay in read-write mode writes the missing provenance in, since it proved which step the entry belongs to.

* test(web-benchmark): complete the provenance of 25 agent recordings

Written by a local read-write replay (no model calls); only recordedFor and createdAt change. The 4 gap entries need a live run with a key.

* test(cache): assert the provenance backfill comes from a replay
…talls pass (tester-army#790)

* fix(e2e): load TypeScript with an oxc loader instead of tsx

* fix(e2e): close loader edge cases: sourceless CommonJS, export maps without extensions, unreadable tsconfig, chained source maps, Windows require paths

* fix(e2e): refuse Bun and Deno runtimes, name oxc load failures, follow Yarn PnP export targets

* perf(e2e): skip the syntax check for files whose only @ is in a package import

* fix(e2e): allow type-only exports in .cts, warn from runner only, Bun check first, docs

* fix(e2e): workers warn about tsconfigs runner did not, FILE: urls, foreign infra errors, review tests
…er-army#775)

* feat(decision): add Clef and Jev executors on the System One choice API

* feat(decision): default Clef transport to clef-flash

* feat(decision): generalize to decision models, address review

* fix(decision): reject non-string decision choices

* fix(decision): skip empty select labels and blank destinations

An empty option label or whitespace-only destination offered as a choice fails at dispatch with INVALID_ARGUMENT and aborts the step instead of letting the model re-aim. Omit them at enumeration. Also qualify the uncertain-action docs: blocking stops the selected action, while earlier confident actions in the step have already run.

* feat(decision): add opt-in vision for Clef, Jev stays text-only

The executor stays text-only by default. With vision true and a transport that declares it, masked viewport pixels ride along as System One images data URLs. Clef declares vision; Jev never receives pixels. Withheld or tainted viewports fall back to semantic text.

* docs(decision): qualify vision cost to calls carrying pixels

* fix(decision): untainted pixel fixtures and honest skill pixels line

The test helper takes a tainted flag so the image-send tests use the only state the runner returns pixels in; tainted captures stay withheld observations. The skill no longer claims pixels are unavailable: screenshots may inform node choices with vision true, without adding point actions.

* bench(testbed): luna vs clef-flash harness and first numbers

Same todo flow under two configs: default agent on gpt-6-luna via OpenCodeX and decisionExecutor on clef-flash. Luna side passed in 23s with 6 model calls; clef-flash pending credentials. Includes the probe script, the docs section, and the flow screenshot.

* docs(decision): fill clef-flash benchmark row

Two no-cache runs both blocked uncertain on the first action (a25 at p=0.619, conf=0.387 against the default 0.9 gates), answering in about a second with 1,543 input tokens.

* docs(decision): benchmark aggregates over 5 runs per side

Luna completed 5/5 flows at ~14s and ~25k input tokens per run; clef-flash blocked uncertain 5/5 at ~1s and ~1.5k tokens per run.

* bench(testbed): parametrized flow, gate sweep, wire-shape finding

Params fix the missing type choices; gates 0.6/0.5 and 0.51/0.3 still block. Direct experiment shows hierarchical structured questions take clef-flash from 0.57 to 0.89 confidence on the same screen.

* feat(decision): hierarchical operation-then-target on structured transports

Structured transports (Clef, Jev) now decide the operation first and the target within it, with object instructions and element records. Flat endpoints keep the single string choice. Live benchmark: first operation 0.909/0.802, first target 0.940/0.775, up from 0.747/0.558 flat.

* docs(decision): compare luna vs clef-flash vs clef with measured cost

* feat(decision): run decisions through AI SDK evaluate, gates opt-in

decisionExecutor now accepts any AI SDK evaluation model that answers
choice questions, such as typeSafeAi.evaluationModel('jev-latest'), and
decides through experimental_evaluate (one choice question per decision,
no retries, so the call budget stays exact). ai is an optional peer that
loads only on that path. evaluationModel() exposes clef(), jev(), and
systemOne() transports to experimental_evaluate, which covers Clef until
it has an AI SDK provider.

The official TypeSafe provider rounds probabilities to two places, so a
long distribution can sum to 0.99: validation now allows half a unit per
rounded option, the AI SDK rule, and jev() declares the two places.

Probability and confidence gates are off by default and stay
configurable. Measuring clef-flash without them surfaced three executor
bugs, each fixed here:
- a lone target is dispatched without a one-option question, which the
  choice API rejects with HTTP 422;
- acting decisions receive the step's action history, and the repeat
  guard counts the same action on the same screen across the whole step,
  so an A-B oscillation blocks instead of burning the call budget;
- the operation and target questions state that typed text is not saved
  until it is submitted.
The structured path also checks the call budget before the target
request (review feedback).

Testbed todo flow, 5 no-cache runs each: clef-flash 5/5 (9.3s mean, 12
calls), clef 5/5 (11.5s, 12 calls), luna unchanged at 5/5 (14.9s).

* bench(testbed): Laya through systemOne

Laya serves the System One API and accepts structured values, so it runs
through systemOne({ structured: true }) unchanged. On the todo flow its
first operation question chose unsupported at 0.29 / 0.10 on four runs,
answered by the multilingual checkpoint its router picked for an English
prompt; the fifth run and further probes hit insufficient_credits.

* fix(decision): address review on rounding, history, and evaluation ids

- validateDecision accepts a rounded distribution only when 1 lies between
  the sums of each value's rounding interval, clipped to [0, 1]; a summed
  per-option tolerance exceeded 1 with many options and admitted two
  options at probability 1.
- Decisions send the 20 most recent history entries, each clipped to 240
  characters, so long steps and large params stay within request limits.
- evaluationModel() builds its answer, confidence, and criteria maps with
  null prototypes, so an id such as __proto__ stays an own entry.
- @ai-sdk/provider moves to dependencies: the public declarations import
  its evaluation types. ai stays an optional peer.

* bench(testbed): pin Laya's English checkpoint

Laya's router sent the full English prompt to its multilingual checkpoint.
The API accepts model 'english' to pin the English one (its routing
metadata reports the explicit model), so the bench config sends it.

* docs(decision): benchmark Laya locally on M1, both checkpoints 0/5

* docs(decision): reconcile laya narrative with run evidence, money not tokens

* feat(decision): caller-overridable transport flags on clef() and jev()

* feat(decision): AI SDK evaluate executor with element table and text model

* fix(decision): empty-tree check and provider pin

* fix(decision): address review on secrets, options, errors, and docs

Answers the cubic threads on the evaluate-only executor: terminal checks reject an incomplete screen instead of judging it, secret markers are stripped at any depth, hidden/disabled options are never offered, select options share the 255-choice cap, select criteria name the option, out-of-range probabilities fail, text-model errors mirror the decision mapping, and the README/docs describe the evaluate-only API.

* fix(decision): count dropped select options and fix README

Counts enabled options inside rows dropped past the element cap as omitted, adds the E2EConfig type import to the README example, and qualifies when type is offered.

* fix(decision): settle terminal claims by the check and match codebase style

The second rejected done/failed claim now ends by the check verdict (fails -> ACTION_FAILED, anything else -> blocked), and the completion check honors the gates. The decision state carries the last 10 actions, check targets show their checked state, checked radios are never toggled, and options dropped by the per-question cap stay out of omitted. The docs example opens /todos in test code since navigate is not offered. Single quotes, template literals, and split long lines across the package.

* fix(decision): show filled fields to verdicts and pin the history window

Verdict views (assert and the completion check) build the element table with typing enabled, so a field the step filled keeps its row and value even when the step itself cannot type. The history test now asserts the window holds the newest ten actions.

* fix(decision): let completion checks see results and name the text key

Page text keeps a named node's text when it differs from the name, as formatNode does, so a status named Greeting that says Welcome back, admin! reaches the verdict. The completion check now asks whether the task is complete and sees the non-secret params and the actions with their typed values; rejected claims stay out. The text prompt names the single text key, asks for no code fence, and prefixes the request so JSON-object response modes accept it.

* fix(decision): keep secret handles out of replayed history and mark check inputs as data

Replayed summaries name a filled secret handle; seeding the history now replaces declared secret names with <secret>, so names reach the model only as secret question criteria. The completion check treats action descriptions, typed values, and params as data, and the guide says the check sees only non-secret values.

* fix(decision): redact only the handle in replayed secret-fill summaries

Matching the declared secret name anywhere in a replayed summary also erased labels and values equal to it, and missed handles the runner cut at 40 characters. Only the quoted handle after fill secret at the start of the summary becomes <secret> now.

* docs(decision): describe the one-request fan-out in the guide intro

* fix(decision): count a scroll that changes what is in view as progress

The fingerprint behind the stall guard hashed node content but not position, so a scroll that only moved the viewport read as page unchanged and the third one blocked the step. It now also hashes whether each node meets the viewport (membership, not coordinates), so a scroll that brings other nodes into view resets the count while one that moves nothing, as at the page bottom, still trips the guard.

* chore: prune @pnpm/exe from the lockfile after merging main

* fix(decision): match the supported Node range in engines

---------

Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
… never ran (tester-army#808)

* test(mobile): fake the clock and shrink PNGs in engine tests, trim duplicates

engine.test.ts 6.4s -> 0.4s: retry backoff and transition budgets on fake timers, small screenshots. Merge or drop tests restating constants or covered elsewhere; prose asserts in errors.test check the headline and recovery only.

* test(mobile): delete the simulator suite no workflow runs

E2E_AGENT_DEVICE_SIMULATOR was never set in CI; the mobile benchmark drives real simulators and emulators. Drops test:simulator and the integration project.

* test(mobile): cover duplicate refind, clipboard step label, unreadable secure screenshot

* test(eas): restore Date.now in afterEach, unique run ids, drop constant asserts

* test(github): drop the factory tautology and markdown asserts core owns

* test(mobile): keep the videoTouches and interactive snapshot option checks

* test(mobile): name the openApp permission refusal by what it checks
…ster-army#807)

* test(agent): fake time for stall, oauth device flows, settle, scroll, typing

Steps and oauth waits ran on real timers: SDK retry backoff, device-flow
intervals, rotation waits, screen settle polls. Same assertions on fake time.
ai-trace overlap test orders tools with a gate instead of timers.

* test(agent): cut duplicate, constant, and removed-surface agent tests

Drops tests pinning defaults and removed keys, permutation rows, and cases
another test already covers. agent-protocol judgment cases live in
agent-judgment-schema; openai-request-shape asserts move into
openai-second-turn. Long config messages assert the key path and remedy.

* test(agent): complete_step keeps the first verdict, blocked needs a blocking code

A turn sending [failed, passed] ended passed if the guard regressed: a false
pass. A blocked verdict without a blockable code must be refused.

* test(agent): fake the retry poll in retryingObserve

* test(agent): bound the stall re-send on fake time, refuse a secret fill into a disabled field or a button

* test(agent): restore viewport direction, whole-screen keyboard note, explicit upper-bound accounting, retry count, maxTurns guard
…ts (tester-army#812)

* test(e2e): run slow expect and locator unit tests on fake time

Assert locator description text in translateLocatorError suffix test.

* test(e2e): cover failing value matchers, serial run status, stackless report

* test(e2e): merge markdown and junit reporter cases, drop restated exact:false test

* test(e2e): replace line-by-line list reporter prose tests with six goldens

Goldens live in tests/fixtures/list-reporter; E2E_GOLDEN_UPDATE=1 rewrites them.
Ports the wide-glyph step label clip from the list-reporter integration test.

* test(e2e): keep the bright badge colors and the finished step clip in list reporter tests
…layed scroll (tester-army#820)

* perf(agent): wait for a lost main list once per folded scroll replay

* chore: changeset for the lost list replay wait

* refactor(agent): one loop for list and viewport scroll repeats
…er-army#819)

A capture that came back empty was retried at each later capture point
(body, settle, finally), so a hung app cost 5 s per point. The engine
call was also awaited without racing the budget, so an engine ignoring
cancellation hung the attempt. Capture once per attempt and abandon the
engine call when the budget ends.
…downloads (tester-army#814)

* fix(secrets): rewrite traces and text downloads whenever the run has a secret value

A registered value a test passed as a plain string (an app.open URL, a fill)
was redacted from titles, labels, and agent input but kept verbatim in
trace.zip: trace and download rewriting was gated on the exposure level,
which only a fill or an engine-held secret raises. Rewriting now depends on
the session ledger holding a value. The engine exposure level had no other
consumer and is removed; exposure now only decides pixels and saved-session
taint.

* fix(secrets): cover a plain-string fill, scope docs to the session's known values

* docs(secrets): download redaction is whole-value; TRACE_WITHHELD outside-path case

* docs(secrets): name the download rewrite failure path
…tester-army#813)

* fix(run): serial group failures reach the exit code and the reporters

- exit code folds serial groups: members carry no attempts, so a group's
  infrastructure error exited 1, and a launch failure (all members skipped) exited 0
- group attempt that fails before any member runs (launch, file load) fails
  the first member with the error, rest skip as serial-predecessor-failed;
  an interrupt there skips every member instead
- a retry that never reached its members keeps the verdicts of the attempt before it
- worker crash during a serial member fails that member with WORKER_CRASH
  instead of skipping it as never started

* docs(test): an interrupted serial launch skips every member
… are (tester-army#806)

redactDownload only matched whole secret values while labeling the file
redaction: "complete", so a cut of a secret (observation-limit style) or a
base64 copy of it could sit in a download the report vouched for as clean.
Trace text entries already go through ledger.redactFragments (whole values,
8-char fragment windows, 4-offset base64/base64url decode), and
docs/security.mdx promises downloads are "rewritten the way trace text is" —
the download path now calls the same method. Unit test covers the fragment
and encoded cases end to end through the artifact sink.
…tester-army#833)

tester-army#774 bumped @e2e-dev/web to 0.12.0. The cache key carries the engine minor,
so every committed entry sat under an old hash and the strict web agent job
failed 25 steps with REPLAY_STALE. Re-recorded live; old entries deleted.
Brings in 49 upstream commits (through 7baf454) so tester-army#772
merges cleanly. Conflicts resolved by keeping upstream and re-adding the
testmu entries: the README package row, the docs workspace devDependency,
and the build and test scripts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@e2e-dev/mobile now pins agent-device 0.21.20, so testmu develops against
the same copy, and its peer floor moves past 0.21.20, which still has no
testmu provider. engines.node matches every other @e2e-dev package after
the oxc loader change. This also drops the lockfile's dangling
agent-device 0.21.18 reference left by the merge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@amankansal-lt
amankansal-lt merged commit 3cc2ef4 into LambdaTest:main Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.