Repository navigation
chore: sync with tester-army/e2e main - #5
Merged
amankansal-lt merged 47 commits intoOct 6, 2026
Merged
Conversation
…tegration proves (tester-army#821) * test(e2e): secret tests inspect model input and inflated trace entries * test(mcp): project tool results and errors pass the ledger; poll instead of sleeping * test(e2e): runProject without an app url runs on the fake engine * test(e2e): secret tests fail their deliberate tap in 500 ms, not the 30 s action timeout * test(e2e): trim agent-act, agent-act-stress, agent-policy duplicates; serial artifact checks move to artifact-store-serial * test(run): runner lifecycle on the fake engine, interrupts on the worker path, cut duplicates * test(web): trim web-platform to unique cases and fold the class and URL-pattern suites into its config-file run * test(e2e): drop integration suites unit and testbed tests already cover * test(e2e): drop a margin tap, wait for the full pipe instead of 2 s, no console noise * test(run): parallel workers on the fake engine with event-driven waits, cover a worker-path timeout * test(e2e): expect.poll scope and expect.soft run on the fake engine * test(run): run events off the browser where no page is read, drop the sink duplicate * test(e2e): trim agent, explore, ai-trace, model-usage duplicates; semantic fallback joins agent-fallback-handoff * test(web): fewer reads against the node replaced every frame * test(e2e): reporters on the fake engine, fold the boot precondition, trim skip-after-failure duplicates * test(e2e): drop trace-cache-speed wall-clock gate; its replay check is in trace-cache-replay * test(e2e): agent-act-verbs pages to row 24, records on the main run, drops duplicate verb and upload checks * test(e2e): trim agent-trace-cache to integration-only proofs, 500 ms on failing expects * test(e2e): fold the stale --strict-cache run into the changed-screen suite * test(e2e): extended fixtures on the fake engine with shorter setup sleeps * test(e2e): trace-modes proves --trace in one run * test(mcp): fold mcp-custom into mcp-server on one server, drop duplicate recording and prose checks * test(e2e): hold four negations instead of seventeen in scripted matchers * test(config): loader-only cases move to config-load, cli-init runs the CLI once per loader shape, trim managed-process duplicates * test(e2e): engine contract stalls one observe, drops cases unit tests own * test(e2e): engine targets keep the hook cases engine-contract lacks, assert the capability code * test(run): Ctrl-C releases what prepare acquired; exit 3, --shard, and CI defaults through the CLI * test(e2e): scripted suites keep what unit tests lack, drop the scene nodes only cut cases used * test(e2e): drop the agent-name and CJK-clip runs unit tests cover * test(e2e): scene root id is internal to the scripted scene * test(e2e): keep the no-model hand-off eviction check, a clean start without a session, and slack on mcp session polls * test(e2e): keep fractional usage, executor model in explore, and --trace off; bound the backpressure wait; never leak APP_URL into fixture runs
…ester-army#804) * test(web): hold each segment's wait for a frame that segment received The observer remembers every page it saw paint, so the wait that closed a segment was satisfied by the first frame that page ever delivered. A segment that painted nothing was then recorded as though it had, and the video assertions fail on it later with `inked` below its threshold. The wait now forgets what earlier segments painted. The added test stops the segment that did paint, opens another, and asserts that closing it runs its timeout out; removing the forget makes it fail. Rebased onto the file as it stands after tester-army#798, which moved the observer inside the test that owns it, so the wait is now the one taking a page rather than reaching for it. Spies are restored and forgotten together, so a failed test cannot leave one behind for the next. * test(web): wait for a painted frame before stopping each mid-attempt video segment The flaking test is the mid-attempt split one, not the per-page one: its first segment sometimes stopped before the screencast delivered a frame, and Playwright writes a white frame then. Share one paint observer between both video tests and drop the test that only checked its own copy. --------- Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com> Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ocess (tester-army#809) * fix(run): catch stray rejections and exceptions in in-process runs An in-process run (rawConfig, e2e explore) installed no unhandledRejection or uncaughtException handling, so a test that left a rejection unhandled crashed the host instead of failing. The in-process runner now routes both the way the worker entry does, for as long as it lives. * fix(run): end an in-process worker on a fatal stray as a process would * fix(run): keep engine disposal errors from a crashed in-process worker * docs(changeset): narrow the in-process fatal rejection case
…aded (tester-army#851) * feat(mcp)!: e2e mcp runs headless by default and takes --headed, like run Sessions were headed outside CI and took --headless, the opposite of e2e run and e2e explore. --headless stays accepted and hidden: it now asks for the default, and a refused flag would reach an MCP host as nothing but a failed connection. * docs(skill): --headed shows the UI when the engine supports it * feat(mcp): open_session takes headed, so the agent picks per session The server's --headed sets the default for sessions that do not say, and headed: true or false overrides it, so a user who asks to watch needs no change to the MCP client config.
…army#789) (tester-army#834) The recorder writes a `type` or `typeText` action with value "" (clearing a field), but the reader rejected it as `invalid-entry`, so the step re-ran the model on every run and failed with REPLAY_STALE under --strict-cache. Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* docs(examples): add standalone Vite project * docs(examples): clarify Vite setup behavior
* docs(examples): add standalone Expo project * fix(examples): correct Expo accessibility and adaptive icon
* docs(examples): add standalone SwiftUI project * fix(examples): announce SwiftUI errors and correct font credit * docs(examples): record SwiftUI verification after review
… runs (tester-army#859) * fix(web): re-resolve a locator that turns ambiguous before its action runs * fix(web): read the dispatch split from call log lines only
…-army#860) * fix(web): evaluate a browser.evaluate string as an expression * docs(web): say evaluate returns a string's value and awaits a promise
…thing in the browsers cache (tester-army#803) * fix(web): provision the chromium build the run launches A first run downloaded the full chromium build, and its own headless launch never needed it, while a machine holding only the full build (`install --no-shell`) read as installed and then failed to launch the shell it needed. The check now answers for the build the run launches, and the run mode reaches the install, so a headed run provisions for its window. Rebased onto the pinned playwright-core: this used `playwrightCliPath` and `require.resolve('playwright/package.json')`, neither of which exists after tester-army#798, so the install of a first run would have failed at runtime. The `PLAYWRIGHT_SKIP_BROWSER_GC=1` the install spawns with is kept, so it still adds what the run needs without collecting the revisions other tools share that cache with, and `e2e-web install` now collects nothing for the same reason. `EnginePrepareInfo.headed` is required, alongside the `headed` an engine already receives on `init`. * test(web): exercise the built-in check and the shell-only install The build-selection tests re-implemented the check instead of calling it, so breaking `isBrowserInstalled` left them green, and nothing asserted `--only-shell` at all. They now run the shipped check from the built package against a real browser cache laid out per build, in a child process because Playwright resolves its cache when it first loads and this file has already loaded it, and they assert `installArgs` for both run modes. * test: hand the run mode to every prepare fixture `EnginePrepareInfo.headed` is required now, so each fixture passes it. * fix(web): platform-correct install test, trim docs and jsdoc --------- Co-authored-by: Feco Linhares <fecolinhares@users.noreply.github.com> Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…#836) * fix(run): redact run-level errors with the run's secrets `serializeError` fell back to no redaction when a call site passed no redactor, and several run-level paths did not: a suite hook failure, engine disposal, a worker's fatal error, an in-memory run's fatal, and the scheduler's own protocol errors. The runner process also never seeded its ledger with the config's static secrets, so an error it serialized before any session opened (a collection failure, a provisioning failure) redacted against nothing. A message that quoted a value the app echoed could reach `report.json` and the GitHub reporter in the clear. The process's live secret ledger (`processSecrets`) is now installed as the default redactor, through a `globalThis` slot so a project's copy of e2e reaches the runner's redactor. `serializeError` uses it when a call site gives none, so a path that forgets its own redactor can no longer leak, while an explicit per-call redactor still wins. The runner seeds that ledger from `config.allSecrets` after resolving the config, the way each worker seeds its own. Verified: new unit and integration tests fail before the change and pass after; the e2e unit suite shows the same failures as the clean tree (pre-existing Windows path baseline only); build, typecheck, and lint pass; the built CLI redacts a collection error that leaked the value on main. * fix(run): trim changeset, cover beforeAll leak, install redactor explicitly --------- Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…ester-army#811) * fix(e2e): keep runner errors from a second module copy classified * test(e2e): pin foreign CollectionError, drop redundant AgentError check
…er-army#863) * docs: add the provider dropdown to the quickstart's API key tab * docs: don't promise a key for providers that read none
…ks again, not MODEL_PROVIDER_FAILED (tester-army#810) The SDK throws ToolChoiceViolationError when a turn forced to call complete_step calls other tools. The loop now handles it like a call to a tool the turn does not offer: the turn counts, each call goes back as an error result with the reply's provider metadata, and the model is asked again on the same turn budget. Out of turns, the step fails as STEP_NO_CONCLUSION.
…sal cases (tester-army#816) * test(cache): run replay backoff on a fake clock trace-replay 141s -> 0.2s, step-cache 14s -> 0.5s. Pin the relocation retry count, and give the end-route test an end wait long enough to enter the wait loop its title claims. * test(cache): drop duplicate route, config, envelope, and helper cases trace-decide restated trace-route through a 4-line wrapper; sameRoute block restated route identity (its unique cases move there). Merge overlapping CI-clamp and envelope tests, drop schema fixture checks schema-fixtures already runs, and cases restating realmSlot, isRecordingMode, attemptSegments, recordedVerdictOf, and empty-state prose. * test(secrets): bound redaction scans by ratio and loose ceilings, not 200ms The linear-pass guard now compares 1x vs 8x input; the near-miss backtracking guard allows 1s. Both still catch quadratic or exponential scans without flaking on a loaded runner. * test(cache): cover verifyEndState's long end wait past the settling backoff * test(cache): pin the trace cache key hash to a golden * test(secrets): refuse an undeclared configured secret and a non-input sink The undeclared case used a name no config held, so the not-configured check refused it too; it now names a configured secret. Sink refusals assert the value is never resolved. * test(sessions): refuse mismatched, expired, and tampered saved sessions * test(cache): keep the options-object mode, minted-id start path, and untainted session load checks * test(internal): keep the function-host realm slot check
…ame masking cases (tester-army#818)
… logs (tester-army#815) * fix(mcp): redact secrets from open_session, close_session, and server logs An engine or app error carrying a configured secret reached the client verbatim in a failed open_session or a close_session Cleanup line, and in the server's log lines. Every fixed tool's result or failure now passes processSecrets, the config's static secrets join it as soon as the config loads, the teardown text passes the session ledger (it is also the stderr shutdown summary), and host log lines are redacted. * fix(mcp): redact shutdown text with every known secret, scope docs claim to loaded configs
…ocess (tester-army#826) * fix(mcp): redact secrets from what user code prints in the e2e mcp process Project tools and the config's top-level code run in the e2e mcp process, and what they wrote to stdout or stderr reached stderr verbatim. Every write now goes through StreamRedactor with processSecrets, as an e2e run worker's output does. Output printed while a config loads is held until its secrets are registered, and withheld if the load fails. * fix(mcp): release unfinished lines per call, keep backpressure, redact uncaught errors * fix(mcp): one redactor for both streams, release tails only with no call in flight, hold uncaught reports during loads * fix(mcp): release a character cut short at shutdown
…and reporters print (tester-army#829)
tester-army#840) Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
…#874) Signed-off-by: dependabot[bot] <support@github.com>
…r-army#877) Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: dependabot[bot] <support@github.com>
…ester-army#879) Signed-off-by: dependabot[bot] <support@github.com>
tester-army#875) Signed-off-by: dependabot[bot] <support@github.com>
…tester-army#880) Signed-off-by: dependabot[bot] <support@github.com>
tester-army#882) Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com> Co-authored-by: Oskar Kwaśniewski <oskar@okwasniewski.com>
Keeps upstream's decision and github entries in AGENTS.md alongside the testmu entry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This brings LambdaTest/e2e
main, which now includes #4'stestmuBrowsers(), up to date with tester-army/e2emain(46 commits). It clears tester-army#772's conflict and puts the browser provider into it.Please merge with "Create a merge commit", not squash or rebase, so
mainfast-forwards and tester-army#772 picks up upstream's history.AGENTS.md. Upstream added thedecisionandgithubentries to the package list, next to ourtestmuentry. All three are kept, and the testmu entry uses upstream's em-dash style.@e2e-dev/webversion (0.12.0, within testmu's>=0.11.0 <1peer), the Node engines, and the DeviceProvider and BrowserProvider contracts are all unchanged.Validation
pnpm install --frozen-lockfileandpnpm checkpass. That covers lint, dead code, typecheck, error codes (84, all documented), peer ranges, install scripts, docs and broken links.@e2e-dev/testmu108/108,@e2e-dev/mobile232/232,@e2e-dev/web504 passed, 1 skipped.git merge-treeagainst tester-armymainis clean.🤖 Generated with Claude Code