Repository navigation
chore: sync with tester-army/e2e main - #3
Merged
amankansal-lt merged 51 commits intoOct 5, 2026
Merged
Conversation
…share a comment (tester-army#720) * fix(github): long marker fields keep a digest so distinct keys never share a comment * docs(changeset): note one-time duplicate for long marker fields
…ent input (tester-army#727) * fix(e2e): redact plain-string secret values in titles, labels, and agent input * fix(e2e): redact explore goal, keep unique() templates aligned with redacted params * test(e2e): use literal escaped marker in junit assertion * fix(e2e): refuse param keys that redact alike, bound explore goal after redaction * fix(e2e): redact titles before NFC, drop duplicate-title heuristic, clarify docs * fix(e2e): redact agent context per step, repair text before cut, re-check title limit * fix(e2e): keep __proto__ param key when redacting, scope params wording to act
…wright does (tester-army#770) * fix(web): list an inline svg as an image, named by its title, as Playwright does Fixes tester-army#749 * test(web): count only compared crosscheck nodes, fix svg fixture comment
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* docs(mobile): document physical devices Closes tester-army#766 * docs(mobile): phone requirements in setup and README * docs(mobile): move physical devices to its own page
…roperty primitives, toHaveURL ignoreCase (tester-army#726) Rebuilt on top of tester-army#579, which already landed matcher option flags, ignoreCase, and unknown-option refusal.
…y#713) * fix(report)!: count interrupted tests apart from failures * fix(report): tell a failed test from its last failed attempt, whatever the cut retry left
…r-army#777) A route segment of letters, a dash, and a digit tail of two or more (`PROJ-016`, `INV-2041`) now abstracts to `:id` like the other minted shapes. These human-readable record ids stayed literals, so a step recorded on one record missed the start-path precondition on the next record and ran live every time. A single trailing digit (`page-2`) stays a route word. Fixes tester-army#769
…r-army#781) * docs(models): add an AI SDK provider picker to the Models page Readers did not see that e2e takes a model from any AI SDK provider. The Models page now has a provider dropdown with 51 providers from ai-sdk.dev, and three setup steps that follow the selection: install, environment, and config. The steps are four instances of one snippet component. Mintlify evaluates each snippet export alone, so the instances share the selection through a store on globalThis and React.useSyncExternalStore. Code renders with Mintlify's CodeBlock for syntax highlighting. * docs(models): an unlisted provider's model needs tool calls and images
* fix(runner): preserve failures when tests skip * refactor(runner): one skipped-after-failure helper, evidence from failing attempt - failureBeforeSkip returns failing attempt's error, evidence, artifacts; list reporter prints those instead of skipped attempt's - someSkippedAfterFailure indexes serial groups once, shared by exit code and run status - collapse skip branch in execute, place failOnSkippedFailure next to retries with JSDoc --------- Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…aunches (tester-army#711) * fix(mobile): never pin an install's file path as the app app.open() launches agent-device 0.21.18 can return an Android install with no package (its binary manifest parser misses aapt2 output, and the aapt fallback needs ANDROID_HOME), echoing the APK path as app. The engine pinned that path, so app.open() failed later with 'Android runtime hints require an installed package name'. installApp now opens by the reported id or the app passed in, and fails with ENGINE_FAILURE when there is neither. app.open() keeps the engine's own UNSUPPORTED_CAPABILITY reason instead of a generic rewrite. * test(mobile): pin the reported package, not the echoed path; clarify installApp failure docs
tester-army#721) * fix(config): refuse secrets and credentials sharing an override env var * docs(environment): name both credential override vars in collision rule
…ion (tester-army#722) * fix(security): allow only http(s) navigation, refuse wrapped schemes on device links view-source:file:// passed the file:/data:/javascript: denylist and loaded local files via app.open, browser.goto, and the navigate verb. resolveNavigationUrl now admits http: and https: only; mobile links also refuse view-source:, blob:, filesystem:. * fix(security): admit exact about:blank, sync rule text in AGENTS.md and SECURITY.md * fix(web): refuse about:blank cookie URL, qualify scheme rule text
…bedded names (tester-army#728) * fix(web): read visibility as rendered, null box, checkable values, embedded names * fix(web): narrow visible queries by the reader's hidden state * fix(web): drop direct text a content-visibility hidden element skips * docs(changeset): note the display contents divergence from Playwright
…ester-army#779) * fix(oauth): reach Copilot models served only over the Responses API copilot() reads the plan's model listing on the first call and picks chat completions or the Responses API for each model, falling back to chat whenever the listing cannot be read so nothing that worked changes. sendCopilotRequest reads the initiator and vision headers from a Responses body (input[]/input_image) as well as a chat one, and the model listing marks an uncallable model 'no chat or responses'. * fix(changeset): schedule a minor release for the Responses endpoint support The changeset asked for a patch while the change adds a call path a model could not reach before; a minor bump is what the release intent states. Addresses the cubic review comment on .changeset/copilot-responses-endpoints.md. * fix(oauth): cancel and never cache a Copilot endpoint lookup that was aborted The first call read the plan's model listing with no signal, and the delegate cached whatever that produced. A listing that stalled therefore left `pending` unresolved and every retry waited on the same promise instead of reaching the model. The signal now reaches the lookup, an aborted lookup rejects instead of being read as a missing endpoint, and a choice that could not be made clears `pending` so the next attempt asks again. A test pins the aborted lookup and the missing-endpoint read as distinct. Addresses the cubic review comment on packages/e2e/src/oauth/copilot.ts. * docs(skill): separate Copilot models served over Responses from uncallable ones The setup reference read as if every non-chat model is reached over the Responses API. Only an enabled model the plan serves solely over `/responses` is; a model served over an API `copilot()` does not speak, or one the plan has not enabled, still cannot be called, and the listing marks those. Addresses the cubic review comment on skills/e2e/references/setup.md. * fix(oauth): pick Copilot endpoint from supported_endpoints alone, retry unreadable listing - drop policy gate: grok-4.5, gpt-5.3-codex serve /responses with no policy - unreadable listing: call over chat, not cached, asked again next call - listing has own 10s timeout; one caller abort no longer fails others - revoked login fails on listing with sign-in hint, no second 401 - missing @ai-sdk/openai names the package to install * fix(init): install @ai-sdk/openai for the Copilot gateway * docs(copilot): unreadable listing retry, init installs both peers --------- Co-authored-by: Hermes Agent <agent@localhost> Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com> Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ywright options (tester-army#723) * fix(web): continue skips earlier routes, decisions take or reject Playwright options continue went through Playwright fallback, so an earlier route answered. It now continues to the network, merging site headers itself; fallback() keeps chaining. continue takes url/method/headers/postData (url via resolveUrl), fulfill takes path/contentType; unknown options are INVALID_ARGUMENT. waitForResponse timeout bounds the match only; setCookies passes the resolved url. * fix(web): route header overrides follow the web({ headers }) grammar * refactor(web): route options reuse rejectUnknownOptions * test(web): route url fake follows the http(s) navigation allowlist * fix(web): fulfill keeps a content-type header over the json or path type
* feat(agent): identify e2e on every model request * docs(changeset): note attribution headers replace provider-level ones * test(agent): check headers on a second act-loop turn
…ts toward negation window (tester-army#725) * fix(expect): unawaited poll fails its own phase, slow first read counts toward negation window - expect.poll is owned by the phase that started it (body, hook, fixture teardown, suite hook); left running it is cancelled and fails that phase with STEP_NOT_AWAITED instead of a later test - pollCondition starts the negation window when the first read that saw it was issued and decides at the deadline without a read past it * fix(expect): own polls by async lineage so a timed-out body never charges the next test * docs(expect): scope STEP_NOT_AWAITED wording to where it surfaces; test reads past body
…tester-army#729) * fix(e2e): redact every observed node field and collapsed cut secrets Observation text, executor trees, and cache anchors leaked a secret held in a test id, frame path, selector, or attribute, and the leading part of a secret with whitespace an engine collapsed before cutting. Nodes now pass one exhaustive redactNode before any reader, and redactCut and redactFragments compare in collapsed-whitespace form. * perf(e2e): skip second redaction pass for already-collapsed fields * perf(e2e): build collapsed reading key in typed arrays, fast path when nothing collapses * refactor(e2e): descriptors, anchors, and relocation read only redacted nodes describeTarget, containerKey, describeNodes, relocateDescriptor, describePosition, describeAnchors, and anchorsPresent take RedactedNode and drop their own weaker redact pass; ReplayHost.redact goes with it. * fix(e2e): re-redact baseline nodes with the current ledger before deriving anchors * test(e2e): count typed-array memory in the large-text redaction guard * fix(e2e): re-redact an action's nodes with the current ledger when describing it * fix(e2e): re-redact an action's container key with the current ledger
…ied debts, kept evidence (tester-army#724) * fix(run): keep --last-failed reruns honest: failure-limit skips, carried debts, kept evidence --last-failed now reruns tests --max-failures skipped. A test or failed suite hook another filter leaves out stays owed in the report's new run.carried until a rerun runs it; the github fold shows carried rows and errors and stays red while anything is carried. A rerun prunes artifacts/ to what the previous report names and writes its attempts under artifacts/rerun-<n>/ instead of wiping the tree and reusing attempt paths. * fix(run): carry a setup's hook debt when no selected test needs it The report has no row for an unneeded setup, so carryForward now asks the collection whether a test still exists instead of the report's rows. * fix(run): reject malformed artifacts in --last-failed reports, tighten rerun docs Distinct ids in the carried fixture row. * fix(run): keep carrying tests of a file a narrowed rerun could not collect * fix(github): a carried interrupted test reads interrupted, not failed, in the folded page
…en (tester-army#732) * fix(cache): stop replays self-finalizing when the effect did not happen Key entries by agent name and a digest of the redacted agent context. Routes carry origin and query, minted values as ids; the end route must match exactly. Anchors record what vanished, control states and values, announcements first; a replay needs evidence the delta happened since the first screen on its end route, and an unseen alert fails it. Unnamed twins need a named row. Steps with no observable delta are not recorded. Policy bumped to conservative/6: every entry re-records. * fix(cache): count call occurrences per agent * fix(cache): hold the end route through the end wait, tighten review nits * fix(cache): fold the agent context for occurrence counts like the key does * refactor(cache): key on the agent context dispatch already redacted * test(e2e): earlier capture redacted again keeps its tree shape
…ne every attempt runs in (tester-army#786) * feat(web): locale and timezoneId options set the language and time zone every attempt runs in * docs(web): timezoneId validation is best effort, the browser has the final word
…rmy#793) device.openLink built its INVALID_ARGUMENT message from the raw input, so a string that fails to parse as an absolute URL still echoed whatever it carried. A malformed magic link is exactly the case where a one-time token sits in that string, and the message reaches logs, the report, and CI output. Name no part of the input, the same rule assertAppId already follows in this file: nothing about a string that failed to parse says which part is safe to repeat, and a token can sit in the query, the path, or the userinfo. linkLabel stays for the step label and the report, where the link did parse. Covered by a unit test over malformed links whose token sits in the query, the path, and the userinfo.
…tester-army#788) * fix(web): a check whose control the click replaces passes and replays check and uncheck click once and read the state back on the same element. A control gone by that read (replaced, or navigated away) took the click, so the action is done instead of NODE_STALE, which the agent turned into LOCATOR_NOT_FOUND and left out of its recording. Fixes ENG-814 * docs(changeset): name the checked radio uncheck refuses
…ter-army#794) * feat(oauth): sign in with OpenCode Console for Zen and Go models * fix(oauth): keep the Console refresh token, isolate tests from OPENCODE_API_KEY A refresh grant without a refresh_token now keeps the previous one, as the ChatGPT and SpaceXAI providers do. The vendor test helper blanks OPENCODE_API_KEY so a key in the shell no longer replaces the stored login under test. Docs and changeset now say the models listing covers Zen and Go only, and that MISCONFIGURED needs a readable workspace config.
…6 cache key (tester-army#796) tester-army#732 bumped REPLAY_POLICY_VERSION to conservative/6 and keyed entries by agent, orphaning every committed entry. Re-recorded tests-agent live and removed the 29 orphaned entries.
…ster-army#798) * feat(web)!: pin playwright-core and add an install command @e2e-dev/web depends on playwright-core 1.63.0 instead of peering on playwright. New e2e-web bin installs browsers for the pinned version: npx @e2e-dev/web install chromium --with-deps. e2e init no longer adds playwright. * test(web): run the e2e-web bin for real, and in CI * test(web): spawn the built e2e-web bin so Node 22.12 runs it
… and install types (tester-army#802)
…ipts }) or browser.addInitScript (tester-army#795) * feat(web): init scripts run before the page's own, configured with web({ initScripts }) or added per test with browser.addInitScript * refactor(web): one owner for configured init scripts, read beside the browser launch * feat(web): export WebInitScript * fix(web): settle the browser launch before a failed init-script read throws; docs examples compile * fix(web): refuse sparse initScripts, keep an own __proto__ key in the init-script argument * docs(web): seed Math.random in the init-script examples * docs(web): trim init-script docs to what users need * docs(skill): addInitScript takes arg with a function only
…anchor (tester-army#799) * fix(cache): record a vanished node only when no end node matches its anchor describeDelta keyed the gone side exactly while deltaHolds matches on the recorded fields alone, so a radio label {text: Express} the pick replaced with a status reading Express was recorded as gone and failed its own check on every replay. * test(web-benchmark): record the radio-replaced-by-summary step
…ecording (tester-army#800) * fix(cache): --strict-cache fails a step whose key changed under its recording A cache key change (REPLAY_POLICY_VERSION, a new key field, an engine minor, the agent's context) turned every committed entry into a no-entry miss, which strict ran live. Strict now lists the file store once per run and fails a no-entry step with REPLAY_STALE when an entry recorded for the same step sits under another key. Entries now record the step's params digest, occurrence, and agent so a repeat or another agent is not mistaken for it. * fix(cache): never take an entry the attempt claimed for a rekeyed step An entry from before the occurrence fields matches by test, target, and instruction; an earlier occurrence's entry is now excluded, so an unrecorded repeat runs live. The message no longer says to delete the old entry, which may still serve another app identity. * fix(cache): match a rekeyed recording only by its whole step An entry without the params, occurrence, and agent could be another call of the same instruction, so strict now ignores it instead of guessing. A confirmed whole replay in read-write mode writes the missing provenance in, since it proved which step the entry belongs to. * test(web-benchmark): complete the provenance of 25 agent recordings Written by a local read-write replay (no model calls); only recordedFor and createdAt change. The 4 gap entries need a live run with a key. * test(cache): assert the provenance backfill comes from a replay
…talls pass (tester-army#790) * fix(e2e): load TypeScript with an oxc loader instead of tsx * fix(e2e): close loader edge cases: sourceless CommonJS, export maps without extensions, unreadable tsconfig, chained source maps, Windows require paths * fix(e2e): refuse Bun and Deno runtimes, name oxc load failures, follow Yarn PnP export targets * perf(e2e): skip the syntax check for files whose only @ is in a package import * fix(e2e): allow type-only exports in .cts, warn from runner only, Bun check first, docs * fix(e2e): workers warn about tsconfigs runner did not, FILE: urls, foreign infra errors, review tests
…er-army#775) * feat(decision): add Clef and Jev executors on the System One choice API * feat(decision): default Clef transport to clef-flash * feat(decision): generalize to decision models, address review * fix(decision): reject non-string decision choices * fix(decision): skip empty select labels and blank destinations An empty option label or whitespace-only destination offered as a choice fails at dispatch with INVALID_ARGUMENT and aborts the step instead of letting the model re-aim. Omit them at enumeration. Also qualify the uncertain-action docs: blocking stops the selected action, while earlier confident actions in the step have already run. * feat(decision): add opt-in vision for Clef, Jev stays text-only The executor stays text-only by default. With vision true and a transport that declares it, masked viewport pixels ride along as System One images data URLs. Clef declares vision; Jev never receives pixels. Withheld or tainted viewports fall back to semantic text. * docs(decision): qualify vision cost to calls carrying pixels * fix(decision): untainted pixel fixtures and honest skill pixels line The test helper takes a tainted flag so the image-send tests use the only state the runner returns pixels in; tainted captures stay withheld observations. The skill no longer claims pixels are unavailable: screenshots may inform node choices with vision true, without adding point actions. * bench(testbed): luna vs clef-flash harness and first numbers Same todo flow under two configs: default agent on gpt-6-luna via OpenCodeX and decisionExecutor on clef-flash. Luna side passed in 23s with 6 model calls; clef-flash pending credentials. Includes the probe script, the docs section, and the flow screenshot. * docs(decision): fill clef-flash benchmark row Two no-cache runs both blocked uncertain on the first action (a25 at p=0.619, conf=0.387 against the default 0.9 gates), answering in about a second with 1,543 input tokens. * docs(decision): benchmark aggregates over 5 runs per side Luna completed 5/5 flows at ~14s and ~25k input tokens per run; clef-flash blocked uncertain 5/5 at ~1s and ~1.5k tokens per run. * bench(testbed): parametrized flow, gate sweep, wire-shape finding Params fix the missing type choices; gates 0.6/0.5 and 0.51/0.3 still block. Direct experiment shows hierarchical structured questions take clef-flash from 0.57 to 0.89 confidence on the same screen. * feat(decision): hierarchical operation-then-target on structured transports Structured transports (Clef, Jev) now decide the operation first and the target within it, with object instructions and element records. Flat endpoints keep the single string choice. Live benchmark: first operation 0.909/0.802, first target 0.940/0.775, up from 0.747/0.558 flat. * docs(decision): compare luna vs clef-flash vs clef with measured cost * feat(decision): run decisions through AI SDK evaluate, gates opt-in decisionExecutor now accepts any AI SDK evaluation model that answers choice questions, such as typeSafeAi.evaluationModel('jev-latest'), and decides through experimental_evaluate (one choice question per decision, no retries, so the call budget stays exact). ai is an optional peer that loads only on that path. evaluationModel() exposes clef(), jev(), and systemOne() transports to experimental_evaluate, which covers Clef until it has an AI SDK provider. The official TypeSafe provider rounds probabilities to two places, so a long distribution can sum to 0.99: validation now allows half a unit per rounded option, the AI SDK rule, and jev() declares the two places. Probability and confidence gates are off by default and stay configurable. Measuring clef-flash without them surfaced three executor bugs, each fixed here: - a lone target is dispatched without a one-option question, which the choice API rejects with HTTP 422; - acting decisions receive the step's action history, and the repeat guard counts the same action on the same screen across the whole step, so an A-B oscillation blocks instead of burning the call budget; - the operation and target questions state that typed text is not saved until it is submitted. The structured path also checks the call budget before the target request (review feedback). Testbed todo flow, 5 no-cache runs each: clef-flash 5/5 (9.3s mean, 12 calls), clef 5/5 (11.5s, 12 calls), luna unchanged at 5/5 (14.9s). * bench(testbed): Laya through systemOne Laya serves the System One API and accepts structured values, so it runs through systemOne({ structured: true }) unchanged. On the todo flow its first operation question chose unsupported at 0.29 / 0.10 on four runs, answered by the multilingual checkpoint its router picked for an English prompt; the fifth run and further probes hit insufficient_credits. * fix(decision): address review on rounding, history, and evaluation ids - validateDecision accepts a rounded distribution only when 1 lies between the sums of each value's rounding interval, clipped to [0, 1]; a summed per-option tolerance exceeded 1 with many options and admitted two options at probability 1. - Decisions send the 20 most recent history entries, each clipped to 240 characters, so long steps and large params stay within request limits. - evaluationModel() builds its answer, confidence, and criteria maps with null prototypes, so an id such as __proto__ stays an own entry. - @ai-sdk/provider moves to dependencies: the public declarations import its evaluation types. ai stays an optional peer. * bench(testbed): pin Laya's English checkpoint Laya's router sent the full English prompt to its multilingual checkpoint. The API accepts model 'english' to pin the English one (its routing metadata reports the explicit model), so the bench config sends it. * docs(decision): benchmark Laya locally on M1, both checkpoints 0/5 * docs(decision): reconcile laya narrative with run evidence, money not tokens * feat(decision): caller-overridable transport flags on clef() and jev() * feat(decision): AI SDK evaluate executor with element table and text model * fix(decision): empty-tree check and provider pin * fix(decision): address review on secrets, options, errors, and docs Answers the cubic threads on the evaluate-only executor: terminal checks reject an incomplete screen instead of judging it, secret markers are stripped at any depth, hidden/disabled options are never offered, select options share the 255-choice cap, select criteria name the option, out-of-range probabilities fail, text-model errors mirror the decision mapping, and the README/docs describe the evaluate-only API. * fix(decision): count dropped select options and fix README Counts enabled options inside rows dropped past the element cap as omitted, adds the E2EConfig type import to the README example, and qualifies when type is offered. * fix(decision): settle terminal claims by the check and match codebase style The second rejected done/failed claim now ends by the check verdict (fails -> ACTION_FAILED, anything else -> blocked), and the completion check honors the gates. The decision state carries the last 10 actions, check targets show their checked state, checked radios are never toggled, and options dropped by the per-question cap stay out of omitted. The docs example opens /todos in test code since navigate is not offered. Single quotes, template literals, and split long lines across the package. * fix(decision): show filled fields to verdicts and pin the history window Verdict views (assert and the completion check) build the element table with typing enabled, so a field the step filled keeps its row and value even when the step itself cannot type. The history test now asserts the window holds the newest ten actions. * fix(decision): let completion checks see results and name the text key Page text keeps a named node's text when it differs from the name, as formatNode does, so a status named Greeting that says Welcome back, admin! reaches the verdict. The completion check now asks whether the task is complete and sees the non-secret params and the actions with their typed values; rejected claims stay out. The text prompt names the single text key, asks for no code fence, and prefixes the request so JSON-object response modes accept it. * fix(decision): keep secret handles out of replayed history and mark check inputs as data Replayed summaries name a filled secret handle; seeding the history now replaces declared secret names with <secret>, so names reach the model only as secret question criteria. The completion check treats action descriptions, typed values, and params as data, and the guide says the check sees only non-secret values. * fix(decision): redact only the handle in replayed secret-fill summaries Matching the declared secret name anywhere in a replayed summary also erased labels and values equal to it, and missed handles the runner cut at 40 characters. Only the quoted handle after fill secret at the start of the summary becomes <secret> now. * docs(decision): describe the one-request fan-out in the guide intro * fix(decision): count a scroll that changes what is in view as progress The fingerprint behind the stall guard hashed node content but not position, so a scroll that only moved the viewport read as page unchanged and the third one blocked the step. It now also hashes whether each node meets the viewport (membership, not coordinates), so a scroll that brings other nodes into view resets the count while one that moves nothing, as at the page bottom, still trips the guard. * chore: prune @pnpm/exe from the lockfile after merging main * fix(decision): match the supported Node range in engines --------- Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
… never ran (tester-army#808) * test(mobile): fake the clock and shrink PNGs in engine tests, trim duplicates engine.test.ts 6.4s -> 0.4s: retry backoff and transition budgets on fake timers, small screenshots. Merge or drop tests restating constants or covered elsewhere; prose asserts in errors.test check the headline and recovery only. * test(mobile): delete the simulator suite no workflow runs E2E_AGENT_DEVICE_SIMULATOR was never set in CI; the mobile benchmark drives real simulators and emulators. Drops test:simulator and the integration project. * test(mobile): cover duplicate refind, clipboard step label, unreadable secure screenshot * test(eas): restore Date.now in afterEach, unique run ids, drop constant asserts * test(github): drop the factory tautology and markdown asserts core owns * test(mobile): keep the videoTouches and interactive snapshot option checks * test(mobile): name the openApp permission refusal by what it checks
…ster-army#807) * test(agent): fake time for stall, oauth device flows, settle, scroll, typing Steps and oauth waits ran on real timers: SDK retry backoff, device-flow intervals, rotation waits, screen settle polls. Same assertions on fake time. ai-trace overlap test orders tools with a gate instead of timers. * test(agent): cut duplicate, constant, and removed-surface agent tests Drops tests pinning defaults and removed keys, permutation rows, and cases another test already covers. agent-protocol judgment cases live in agent-judgment-schema; openai-request-shape asserts move into openai-second-turn. Long config messages assert the key path and remedy. * test(agent): complete_step keeps the first verdict, blocked needs a blocking code A turn sending [failed, passed] ended passed if the guard regressed: a false pass. A blocked verdict without a blockable code must be refused. * test(agent): fake the retry poll in retryingObserve * test(agent): bound the stall re-send on fake time, refuse a secret fill into a disabled field or a button * test(agent): restore viewport direction, whole-screen keyboard note, explicit upper-bound accounting, retry count, maxTurns guard
…ts (tester-army#812) * test(e2e): run slow expect and locator unit tests on fake time Assert locator description text in translateLocatorError suffix test. * test(e2e): cover failing value matchers, serial run status, stackless report * test(e2e): merge markdown and junit reporter cases, drop restated exact:false test * test(e2e): replace line-by-line list reporter prose tests with six goldens Goldens live in tests/fixtures/list-reporter; E2E_GOLDEN_UPDATE=1 rewrites them. Ports the wide-glyph step label clip from the list-reporter integration test. * test(e2e): keep the bright badge colors and the finished step clip in list reporter tests
…layed scroll (tester-army#820) * perf(agent): wait for a lost main list once per folded scroll replay * chore: changeset for the lost list replay wait * refactor(agent): one loop for list and viewport scroll repeats
…er-army#819) A capture that came back empty was retried at each later capture point (body, settle, finally), so a hung app cost 5 s per point. The engine call was also awaited without racing the budget, so an engine ignoring cancellation hung the attempt. Capture once per attempt and abandon the engine call when the budget ends.
…downloads (tester-army#814) * fix(secrets): rewrite traces and text downloads whenever the run has a secret value A registered value a test passed as a plain string (an app.open URL, a fill) was redacted from titles, labels, and agent input but kept verbatim in trace.zip: trace and download rewriting was gated on the exposure level, which only a fill or an engine-held secret raises. Rewriting now depends on the session ledger holding a value. The engine exposure level had no other consumer and is removed; exposure now only decides pixels and saved-session taint. * fix(secrets): cover a plain-string fill, scope docs to the session's known values * docs(secrets): download redaction is whole-value; TRACE_WITHHELD outside-path case * docs(secrets): name the download rewrite failure path
…tester-army#813) * fix(run): serial group failures reach the exit code and the reporters - exit code folds serial groups: members carry no attempts, so a group's infrastructure error exited 1, and a launch failure (all members skipped) exited 0 - group attempt that fails before any member runs (launch, file load) fails the first member with the error, rest skip as serial-predecessor-failed; an interrupt there skips every member instead - a retry that never reached its members keeps the verdicts of the attempt before it - worker crash during a serial member fails that member with WORKER_CRASH instead of skipping it as never started * docs(test): an interrupted serial launch skips every member
… are (tester-army#806) redactDownload only matched whole secret values while labeling the file redaction: "complete", so a cut of a secret (observation-limit style) or a base64 copy of it could sit in a download the report vouched for as clean. Trace text entries already go through ledger.redactFragments (whole values, 8-char fragment windows, 4-offset base64/base64url decode), and docs/security.mdx promises downloads are "rewritten the way trace text is" — the download path now calls the same method. Unit test covers the fragment and encoded cases end to end through the artifact sink.
…tester-army#833) tester-army#774 bumped @e2e-dev/web to 0.12.0. The cache key carries the engine minor, so every committed entry sat under an old hash and the strict web agent job failed 25 steps with REPLAY_STALE. Re-recorded live; old entries deleted.
Brings in 49 upstream commits (through 7baf454) so tester-army#772 merges cleanly. Conflicts resolved by keeping upstream and re-adding the testmu entries: the README package row, the docs workspace devDependency, and the build and test scripts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@e2e-dev/mobile now pins agent-device 0.21.20, so testmu develops against the same copy, and its peer floor moves past 0.21.20, which still has no testmu provider. engines.node matches every other @e2e-dev package after the oxc loader change. This also drops the lockfile's dangling agent-device 0.21.18 reference left by the merge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Brings LambdaTest/e2e
mainup to date with tester-army/e2emain(49 commits) so that tester-army#772, whose head is this fork'smain, stops conflicting.Please merge this with a merge commit, not squash or rebase, so
mainfast-forwards and tester-army#772 picks up upstream's history.8773058bmerges tester-armymain. Conflicts were inREADME.md,docs/package.jsonandpackage.json. In each, upstream's version is kept and only the testmu additions are re-added: the README row, the docs devDependency, and the build/test scripts.4a6def63brings@e2e-dev/testmuin line with upstream:@e2e-dev/mobile's pin;>0.21.20 <1;^22.22.3 || >=24.8.0;No testmu source changes were needed: upstream didn't touch the DeviceProvider/
record()contract or the docs structure testmu uses.Validation
Run on Node 22.22.3:
lint,check:dead-code,typecheck,docs:check-errors,check:peer-ranges,check:install-scripts,docs:checkandtestall pass. testmu: 73 tests; mobile: 232; e2e: 3160.git merge-treeagainst tester-armymainis clean.🤖 Generated with Claude Code