Skip to content

chore: sync with tester-army/e2e main - #5

Merged
amankansal-lt merged 47 commits into
LambdaTest:mainfrom
amankansal-lt:chore/sync-upstream-main-2
Oct 6, 2026
Merged

amankansal-lt merged 47 commits into
LambdaTest:mainfrom
amankansal-lt:chore/sync-upstream-main-2

Conversation

@amankansal-lt

Copy link
Copy Markdown
Collaborator

Summary

This brings LambdaTest/e2e main, which now includes #4's testmuBrowsers(), up to date with tester-army/e2e main (46 commits). It clears tester-army#772's conflict and puts the browser provider into it.

Please merge with "Create a merge commit", not squash or rebase, so main fast-forwards and tester-army#772 picks up upstream's history.

  • The only conflict was AGENTS.md. Upstream added the decision and github entries to the package list, next to our testmu entry. All three are kept, and the testmu entry uses upstream's em-dash style.
  • Nothing testmu depends on moved upstream. The agent-device pin (0.21.20), the @e2e-dev/web version (0.12.0, within testmu's >=0.11.0 <1 peer), the Node engines, and the DeviceProvider and BrowserProvider contracts are all unchanged.

Validation

  • pnpm install --frozen-lockfile and pnpm check pass. That covers lint, dead code, typecheck, error codes (84, all documented), peer ranges, install scripts, docs and broken links.
  • Tests: @e2e-dev/testmu 108/108, @e2e-dev/mobile 232/232, @e2e-dev/web 504 passed, 1 skipped.
  • git merge-tree against tester-army main is clean.

🤖 Generated with Claude Code

okwasniewski and others added 30 commits October 5, 2026 10:40
…tegration proves (tester-army#821)

* test(e2e): secret tests inspect model input and inflated trace entries

* test(mcp): project tool results and errors pass the ledger; poll instead of sleeping

* test(e2e): runProject without an app url runs on the fake engine

* test(e2e): secret tests fail their deliberate tap in 500 ms, not the 30 s action timeout

* test(e2e): trim agent-act, agent-act-stress, agent-policy duplicates; serial artifact checks move to artifact-store-serial

* test(run): runner lifecycle on the fake engine, interrupts on the worker path, cut duplicates

* test(web): trim web-platform to unique cases and fold the class and URL-pattern suites into its config-file run

* test(e2e): drop integration suites unit and testbed tests already cover

* test(e2e): drop a margin tap, wait for the full pipe instead of 2 s, no console noise

* test(run): parallel workers on the fake engine with event-driven waits, cover a worker-path timeout

* test(e2e): expect.poll scope and expect.soft run on the fake engine

* test(run): run events off the browser where no page is read, drop the sink duplicate

* test(e2e): trim agent, explore, ai-trace, model-usage duplicates; semantic fallback joins agent-fallback-handoff

* test(web): fewer reads against the node replaced every frame

* test(e2e): reporters on the fake engine, fold the boot precondition, trim skip-after-failure duplicates

* test(e2e): drop trace-cache-speed wall-clock gate; its replay check is in trace-cache-replay

* test(e2e): agent-act-verbs pages to row 24, records on the main run, drops duplicate verb and upload checks

* test(e2e): trim agent-trace-cache to integration-only proofs, 500 ms on failing expects

* test(e2e): fold the stale --strict-cache run into the changed-screen suite

* test(e2e): extended fixtures on the fake engine with shorter setup sleeps

* test(e2e): trace-modes proves --trace in one run

* test(mcp): fold mcp-custom into mcp-server on one server, drop duplicate recording and prose checks

* test(e2e): hold four negations instead of seventeen in scripted matchers

* test(config): loader-only cases move to config-load, cli-init runs the CLI once per loader shape, trim managed-process duplicates

* test(e2e): engine contract stalls one observe, drops cases unit tests own

* test(e2e): engine targets keep the hook cases engine-contract lacks, assert the capability code

* test(run): Ctrl-C releases what prepare acquired; exit 3, --shard, and CI defaults through the CLI

* test(e2e): scripted suites keep what unit tests lack, drop the scene nodes only cut cases used

* test(e2e): drop the agent-name and CJK-clip runs unit tests cover

* test(e2e): scene root id is internal to the scripted scene

* test(e2e): keep the no-model hand-off eviction check, a clean start without a session, and slack on mcp session polls

* test(e2e): keep fractional usage, executor model in explore, and --trace off; bound the backpressure wait; never leak APP_URL into fixture runs
…ester-army#804)

* test(web): hold each segment's wait for a frame that segment received

The observer remembers every page it saw paint, so the wait that closed a
segment was satisfied by the first frame that page ever delivered. A segment
that painted nothing was then recorded as though it had, and the video
assertions fail on it later with `inked` below its threshold.

The wait now forgets what earlier segments painted. The added test stops the
segment that did paint, opens another, and asserts that closing it runs its
timeout out; removing the forget makes it fail.

Rebased onto the file as it stands after tester-army#798, which moved the observer
inside the test that owns it, so the wait is now the one taking a page rather
than reaching for it. Spies are restored and forgotten together, so a failed
test cannot leave one behind for the next.

* test(web): wait for a painted frame before stopping each mid-attempt video segment

The flaking test is the mid-attempt split one, not the per-page one: its
first segment sometimes stopped before the screencast delivered a frame,
and Playwright writes a white frame then. Share one paint observer between
both video tests and drop the test that only checked its own copy.

---------

Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com>
Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ocess (tester-army#809)

* fix(run): catch stray rejections and exceptions in in-process runs

An in-process run (rawConfig, e2e explore) installed no unhandledRejection or
uncaughtException handling, so a test that left a rejection unhandled crashed
the host instead of failing. The in-process runner now routes both the way the
worker entry does, for as long as it lives.

* fix(run): end an in-process worker on a fatal stray as a process would

* fix(run): keep engine disposal errors from a crashed in-process worker

* docs(changeset): narrow the in-process fatal rejection case
…aded (tester-army#851)

* feat(mcp)!: e2e mcp runs headless by default and takes --headed, like run

Sessions were headed outside CI and took --headless, the opposite of
e2e run and e2e explore. --headless stays accepted and hidden: it now
asks for the default, and a refused flag would reach an MCP host as
nothing but a failed connection.

* docs(skill): --headed shows the UI when the engine supports it

* feat(mcp): open_session takes headed, so the agent picks per session

The server's --headed sets the default for sessions that do not say,
and headed: true or false overrides it, so a user who asks to watch
needs no change to the MCP client config.
…army#789) (tester-army#834)

The recorder writes a `type` or `typeText` action with value "" (clearing a
field), but the reader rejected it as `invalid-entry`, so the step re-ran the
model on every run and failed with REPLAY_STALE under --strict-cache.

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* docs(examples): add standalone Vite project

* docs(examples): clarify Vite setup behavior
* docs(examples): add standalone Expo project

* fix(examples): correct Expo accessibility and adaptive icon
* docs(examples): add standalone SwiftUI project

* fix(examples): announce SwiftUI errors and correct font credit

* docs(examples): record SwiftUI verification after review
… runs (tester-army#859)

* fix(web): re-resolve a locator that turns ambiguous before its action runs

* fix(web): read the dispatch split from call log lines only
…-army#860)

* fix(web): evaluate a browser.evaluate string as an expression

* docs(web): say evaluate returns a string's value and awaits a promise
…thing in the browsers cache (tester-army#803)

* fix(web): provision the chromium build the run launches

A first run downloaded the full chromium build, and its own headless launch
never needed it, while a machine holding only the full build (`install
--no-shell`) read as installed and then failed to launch the shell it
needed. The check now answers for the build the run launches, and the run
mode reaches the install, so a headed run provisions for its window.

Rebased onto the pinned playwright-core: this used `playwrightCliPath` and
`require.resolve('playwright/package.json')`, neither of which exists after
tester-army#798, so the install of a first run would have failed at runtime. The
`PLAYWRIGHT_SKIP_BROWSER_GC=1` the install spawns with is kept, so it still
adds what the run needs without collecting the revisions other tools share
that cache with, and `e2e-web install` now collects nothing for the same
reason.

`EnginePrepareInfo.headed` is required, alongside the `headed` an engine
already receives on `init`.

* test(web): exercise the built-in check and the shell-only install

The build-selection tests re-implemented the check instead of calling it,
so breaking `isBrowserInstalled` left them green, and nothing asserted
`--only-shell` at all. They now run the shipped check from the built
package against a real browser cache laid out per build, in a child process
because Playwright resolves its cache when it first loads and this file has
already loaded it, and they assert `installArgs` for both run modes.

* test: hand the run mode to every prepare fixture

`EnginePrepareInfo.headed` is required now, so each fixture passes it.

* fix(web): platform-correct install test, trim docs and jsdoc

---------

Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com>
Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…#836)

* fix(run): redact run-level errors with the run's secrets

`serializeError` fell back to no redaction when a call site passed no
redactor, and several run-level paths did not: a suite hook failure, engine
disposal, a worker's fatal error, an in-memory run's fatal, and the
scheduler's own protocol errors. The runner process also never seeded its
ledger with the config's static secrets, so an error it serialized before
any session opened (a collection failure, a provisioning failure) redacted
against nothing. A message that quoted a value the app echoed could reach
`report.json` and the GitHub reporter in the clear.

The process's live secret ledger (`processSecrets`) is now installed as the
default redactor, through a `globalThis` slot so a project's copy of e2e
reaches the runner's redactor. `serializeError` uses it when a call site
gives none, so a path that forgets its own redactor can no longer leak,
while an explicit per-call redactor still wins. The runner seeds that ledger
from `config.allSecrets` after resolving the config, the way each worker
seeds its own.

Verified: new unit and integration tests fail before the change and pass
after; the e2e unit suite shows the same failures as the clean tree
(pre-existing Windows path baseline only); build, typecheck, and lint pass;
the built CLI redacts a collection error that leaked the value on main.

* fix(run): trim changeset, cover beforeAll leak, install redactor explicitly

---------

Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ester-army#811)

* fix(e2e): keep runner errors from a second module copy classified

* test(e2e): pin foreign CollectionError, drop redundant AgentError check
…er-army#863)

* docs: add the provider dropdown to the quickstart's API key tab

* docs: don't promise a key for providers that read none
…ks again, not MODEL_PROVIDER_FAILED (tester-army#810)

The SDK throws ToolChoiceViolationError when a turn forced to call complete_step calls other tools. The loop now handles it like a call to a tool the turn does not offer: the turn counts, each call goes back as an error result with the reply's provider metadata, and the model is asked again on the same turn budget. Out of turns, the step fails as STEP_NO_CONCLUSION.
…sal cases (tester-army#816)

* test(cache): run replay backoff on a fake clock

trace-replay 141s -> 0.2s, step-cache 14s -> 0.5s. Pin the relocation
retry count, and give the end-route test an end wait long enough to enter
the wait loop its title claims.

* test(cache): drop duplicate route, config, envelope, and helper cases

trace-decide restated trace-route through a 4-line wrapper; sameRoute block
restated route identity (its unique cases move there). Merge overlapping
CI-clamp and envelope tests, drop schema fixture checks schema-fixtures
already runs, and cases restating realmSlot, isRecordingMode,
attemptSegments, recordedVerdictOf, and empty-state prose.

* test(secrets): bound redaction scans by ratio and loose ceilings, not 200ms

The linear-pass guard now compares 1x vs 8x input; the near-miss
backtracking guard allows 1s. Both still catch quadratic or exponential
scans without flaking on a loaded runner.

* test(cache): cover verifyEndState's long end wait past the settling backoff

* test(cache): pin the trace cache key hash to a golden

* test(secrets): refuse an undeclared configured secret and a non-input sink

The undeclared case used a name no config held, so the not-configured
check refused it too; it now names a configured secret. Sink refusals
assert the value is never resolved.

* test(sessions): refuse mismatched, expired, and tampered saved sessions

* test(cache): keep the options-object mode, minted-id start path, and untainted session load checks

* test(internal): keep the function-host realm slot check
… logs (tester-army#815)

* fix(mcp): redact secrets from open_session, close_session, and server logs

An engine or app error carrying a configured secret reached the client
verbatim in a failed open_session or a close_session Cleanup line, and in
the server's log lines. Every fixed tool's result or failure now passes
processSecrets, the config's static secrets join it as soon as the config
loads, the teardown text passes the session ledger (it is also the stderr
shutdown summary), and host log lines are redacted.

* fix(mcp): redact shutdown text with every known secret, scope docs claim to loaded configs
…ocess (tester-army#826)

* fix(mcp): redact secrets from what user code prints in the e2e mcp process

Project tools and the config's top-level code run in the e2e mcp process,
and what they wrote to stdout or stderr reached stderr verbatim. Every
write now goes through StreamRedactor with processSecrets, as an e2e run
worker's output does. Output printed while a config loads is held until
its secrets are registered, and withheld if the load fails.

* fix(mcp): release unfinished lines per call, keep backpressure, redact uncaught errors

* fix(mcp): one redactor for both streams, release tails only with no call in flight, hold uncaught reports during loads

* fix(mcp): release a character cut short at shutdown
okwasniewski and others added 17 commits October 5, 2026 16:22
tester-army#840)

Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…#874)

Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: dependabot[bot] <support@github.com>
tester-army#882)

Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
Keeps upstream's decision and github entries in AGENTS.md alongside the
testmu entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@amankansal-lt
amankansal-lt merged commit 24a10a4 into LambdaTest:main Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.