Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
14 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -4,5 +4,11 @@ TEXT_MODEL_API_KEY=
TEXT_MODEL_BASE_URL=https://openrouter.ai/api/v1
TEXT_MODEL=inception/mercury-2.5
TEXT_MODEL_REASONING=none
# Optional. Keep this outside the repository. Each home is a separate session.
# Optional vision: also set visionEnabled=true in the selected home's config.json.
# Screenshots go to this explicitly configured provider. No default vision model.
MIDSCENE_MODEL_API_KEY=
MIDSCENE_MODEL_BASE_URL=
MIDSCENE_MODEL_NAME=
MIDSCENE_MODEL_FAMILY=
# Each home/profile is a separate session; keep it outside repositories.
# FASTEST_E2E_HOME=/absolute/path/to/automation-state
30 changes: 22 additions & 8 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ permissions:
jobs:
check:
runs-on: ubuntu-latest
timeout-minutes: 15
timeout-minutes: 20
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
Expand All @@ -23,11 +23,25 @@ jobs:
run: |
npm ci --ignore-scripts
uv sync --project worker --locked
- name: Typecheck and test CLI, MCP, and worker
- name: Audit runtime dependencies
run: npm audit --omit=dev --audit-level=high
- name: Typecheck and unit/protocol tests
if: always() && !cancelled()
run: npm run check
- name: Real Chrome smoke test, without model API calls
if: ${{ !cancelled() }}
run: uv run --project worker --no-sync python worker/tests/browser_smoke.py
- name: Public CLI lifecycle, without model API calls
if: ${{ !cancelled() }}
run: node test/lifecycle.smoke.mjs
- name: Legacy upstream browser and CLI lifecycle
if: always() && !cancelled()
run: |
uv run --project worker --no-sync python worker/tests/browser_smoke.py
node test/lifecycle.smoke.mjs
- name: Managed upstream Jev and crash receipts
if: always() && !cancelled()
run: uv run --project worker --no-sync python worker/tests/managed_smoke.py
- name: Scoped adapter and native MCP images
if: always() && !cancelled()
run: timeout 90s node test/adapter.smoke.mjs
- name: Public workflows and real crash recovery
if: always() && !cancelled()
run: timeout 180s node test/workflows.smoke.mjs
- name: Actual Midscene with controlled provider
if: always() && !cancelled()
run: timeout 90s node test/vision.smoke.mjs
12 changes: 5 additions & 7 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,11 @@
# fastest-e2e

Read `package.json` for commands and `README.md` for supported behavior. Run `npm run check` before committing. Browser integration tests are documented in the README.
Read `package.json` for commands, `README.md` for usage, and `docs/design.md` for runtime/recovery invariants. Run `npm run check`; browser suites are in CI and README. No parallel design/process documents.

For planned runtime changes, read [the design and checklist](docs/design.md). Keep proposals out of operational skills until implemented. A saved run is not saved browser state; distinguish validated continuation from reconstruction, restart, and unrecoverable loss. Keep design decisions in that one document, not parallel process documents.
CLI and MCP use the same Effect v4 handlers and schemas. Reuse upstream Jev, Playwright, and Midscene. Do not build another agent loop. Pin dependencies with both lockfiles and Actions with full commit SHAs.

CLI and MCP call the same Effect v4 service. Keep the Jev action loop upstream; `worker/bridge.py` owns only session handoff, assertions, and the process protocol. Pin upgrades as one reviewed change with both lock files updated.
A saved run is not saved browser state. Preserve profile/target/document ownership, dispatch uncertainty, cumulative budgets, and failed attempts. Never retry uncertain production writes or use reconstruction to hide a failed persistence test. Missing evidence cannot become a pass.

Browser selection must fail closed. Never add automatic personal-profile discovery, implicit cloud fallback, or retries of uncertain production mutations. Keep stdout valid JSON or MCP; redact credentials and avoid raw model error dumps.
Keep stdout JSON/MCP, errors redacted, provider use explicit, and recovery inputs allowlisted. Literal task inputs are persisted; sensitive values need local environment references. Do not claim lossless restore, complete redaction, or an OS sandbox.

A test pass requires fresh explicit evidence. A model's completion claim is not evidence. Preserve blocked/failed tabs and failed attempts for diagnosis.

Keep the three skills short. Put setup details in `docs/setup.md`, fallback details in `docs/fallback.md`, and command options in `--help`. Change shared behavior in one place.
Keep the three skills short. Put setup in `docs/setup.md`, tests in `docs/testing.md`, and fallback in `docs/fallback.md`. Change shared behavior in one place and test CLI/MCP parity.
108 changes: 51 additions & 57 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,21 +1,23 @@
# fastest-e2e

Use a logged-in browser or test a deployed feature from your coding agent or terminal. Application setup and cleanup use the UI. No application API or database access is required.
Browser use and evidence-based E2E tests from a coding agent, CLI, or MCP. Use a dedicated, logged-in Chrome profile, normally headless. Application setup and cleanup stay in the UI.

```mermaid
flowchart LR
Caller["You or your coding agent"] --> Entry["CLI / MCP"]
Entry --> Jev["Jev + Mercury"]
Jev --> Harness["Browser Harness"]
Entry -. "same-tab fallback" .-> Harness
Harness --> Chrome["Dedicated Chrome profile"]
Agent["You / coding agent"] --> Run["Effect CLI + MCP / durable run"]
Run --> Jev["Jev: autonomous DOM tasks"]
Run --> Steps["Playwright: scoped UI steps"]
Run --> Vision["Midscene: visual workflows"]
Jev --> Chrome["Same dedicated Chrome"]
Steps --> Chrome
Vision --> Chrome
Chrome --> Evidence["Checks + extraction + images"]
Evidence --> Run
```

The CLI and MCP share an Effect v4 runtime. The Python worker reuses upstream Jev Ultrafast and Browser Harness. The profile contains only the accounts you sign into there, not your personal Chrome sessions.

## First run

Requires Node 24.18 or newer, uv, and Google Chrome or Chromium. There is no published npm package yet.
Requires Node 24.18+, uv, and installed Chrome or Chromium. No npm package is published yet.

```sh
git clone https://github.com/BleedingDev/fastest-e2e.git
Expand All @@ -28,47 +30,37 @@ fastest-e2e install
fastest-e2e start --headed
```

Sign into only the intended accounts in the new Chrome window. Keep Chrome sync off. `init` uses `~/.fastest-e2e`; [setup](docs/setup.md) covers custom profiles and installation troubleshooting.

Make `TYPESAFE_API_KEY` and `TEXT_MODEL_API_KEY` available in the invoking shell or secret manager. The default text provider is OpenRouter. See [.env.example](.env.example) and [environment setup](docs/setup.md) for file-based loading. Keep keys out of chat and git. Page and field context goes to the configured model providers.

Then switch the same profile to headless:
Sign into only the intended accounts; keep Chrome sync off. Supply `TYPESAFE_API_KEY` and `TEXT_MODEL_API_KEY` through your invoking environment. Then:

```sh
fastest-e2e stop
fastest-e2e start
fastest-e2e doctor
```

In the JSON, `configured`, `workerInstalled`, `connected`, `jevKeyPresent`, and `textKeyPresent` should all be `true`. `doctor` does not validate keys with providers, and its exit code alone does not check key presence. Headed and headless browsers cannot use the directory simultaneously. Logins may expire.
Inspect the readiness fields, not just the exit code. Key presence does not prove validity. [Setup](docs/setup.md) covers custom profiles, provider configuration, and registering the three skills.

## Use the browser

Replace the URL with your application:

```sh
fastest-e2e run --url https://your-app.example/settings \
"Read the current display name. Leave all settings unchanged."
"Open account settings. Leave values unchanged."
fastest-e2e inspect --run RUN_ID --view page
fastest-e2e screenshot --run RUN_ID
fastest-e2e close --run RUN_ID
```

The command returns a status and `targetId`, not the extracted answer. `done` means Jev stopped, not that the requested outcome was independently verified. Read the page and check the result:

```sh
fastest-e2e inspect --target TARGET_ID
fastest-e2e close --target TARGET_ID
```
Replace `RUN_ID` with the returned `runId`. Summaries include progress, attempt history, evidence references, and budgets. Page inspection returns focused controls/text; screenshots return a local image path. MCP `browser_screenshot` returns the image itself, not a path-only response.

Replace `TARGET_ID` with the returned identifier. `inspect` returns the URL, title, and page text. For a control value, a screenshot, or an unsupported interaction, use the [same-tab fallback](docs/fallback.md).
For one-call structured reading, edit [read.task.json](examples/read.task.json), then run `fastest-e2e run --file read.task.json`. The result's `extraction` contains values and evidence, or explicit field errors. No model call is needed for its deterministic checkpoint and DOM extraction.

## Test a PR or feature

After [registering the skills](docs/setup.md#skills), give your coding agent a concrete request:

> Use browser-test to test PR #123 at https://preview.example/settings with the account already logged in. Confirm the deployment contains this PR. Keep this run read-only and report checks that need permission to change data.
Ask an agent with the skills installed:

The agent reads the change, chooses checks, runs them, and reports evidence. The CLI itself does not fetch PRs or deploy code. A production URL that lacks the PR cannot validate it.
> Use browser-test to test PR #123 at the deployed preview URL with the logged-in test account. Establish which build is deployed, define checks before acting, and report evidence and untested behavior.

To run a test yourself, copy [account.test.json](examples/account.test.json) and replace its URL and expected heading:
The coding agent reads the PR; the browser runtime does not fetch or deploy it. To run a test yourself, adapt [account.test.json](examples/account.test.json):

```json
{
Expand All @@ -85,55 +77,57 @@ To run a test yourself, copy [account.test.json](examples/account.test.json) and
fastest-e2e test --file account.test.json
```

This is a read-only goal, not an enforced read-only browser mode. For authorized write tests, adapt [settings.test.json](examples/settings.test.json) and agree on cleanup. [Testing reference](docs/testing.md) explains assertion types and persistence checks.
A read-only goal is guidance, not an enforced read-only browser. Authorize production changes before running them. [Testing](docs/testing.md) covers explicit UI steps, frame/shadow scopes, checkpoints, extraction, and cleanup.

| Result | Meaning | Exit code |
| Result | Meaning | Exit |
| --- | --- | --- |
| `done` | Browser task ended without independent assertions. | 0 |
| `passed` | Jev finished and every explicit check passed. | 0 |
| `failed` | An explicit check did not match. | 1 |
| `blocked` | Execution or verification could not finish. | 2 |
| `done` | Execution completed without a test verdict. | 0 |
| `passed` | Every declared test check passed. | 0 |
| `failed` | A declared check did not match. | 1 |
| `blocked` | Execution, extraction, or verification could not finish. | 2 |

Failed attempts are retained. `verify --run RUN_ID` rechecks saved assertions without replaying actions or replacing the original verdict. Successful tests close their task tab unless `keepTab: true`.

Results include `checks` with expected and actual values. Failed and blocked tests retain their tab. Inspect it before retrying a mutation. A successful test closes its tab unless `keepTab` is `true`.
## Choose execution, not a different browser

## Use your coding agent
Jev remains the default. Explicit `steps` use Playwright without model calls; a `goal` step delegates only that subtask. `frames` and `shadow: "open"` support scoped controls and checks. Closed-root DOM access returns unsupported rather than an accidental absence/pass.

| Skill | Ask it to do |
| --- | --- |
| [setup-fastest-e2e](skills/setup-fastest-e2e/SKILL.md) | Install or repair the runtime and register the skills. |
| [browser-use](skills/browser-use/SKILL.md) | Complete a browser task and verify what happened. |
| [browser-test](skills/browser-test/SKILL.md) | Test a PR, feature, fix, or regression against a specific deployment. |
For image/canvas workflows, configure vision once, then use `--engine vision`. `--engine auto` permits same-tab Jev-to-Midscene handoff only when the current autonomous segment stopped before dispatching any action. After partial autonomous work, inspect and explicitly resume with the remaining goal. Failed assertions, uncertain submissions, and cancellation never trigger automatic replay. [Fallback](docs/fallback.md) gives the commands.

Start with the setup skill. [Registration and MCP configuration](docs/setup.md#skills) cover using the same installation from different agents. `fastest-e2e mcp` exposes high-level tools over stdio. Trusted Python fallback is CLI-accessible; MCP exposes it only with `mcp --allow-scripts`.
## Recover work, not imaginary browser state

## Sessions and limits
```sh
fastest-e2e inspect --run RUN_ID
fastest-e2e inspect --run RUN_ID --view page
fastest-e2e resume --run RUN_ID --revision REVISION
```

Chrome and Browser Harness stay warm between tasks. Each `run` or `test` opens a new owned tab with the profile's stored authentication; it does not resume an existing tab. Use the returned target for inspection or fallback. Concurrent operations on one configured home are rejected rather than queued. Separate homes **and** profiles are required for independent sessions.
Use the latest returned revision. Resume verifies the live document and state. A reload or crash may make continuation impossible. An explicitly declared reconstruction plan can rebuild allowed fields through the UI; restart requires a declared repeat-safe workflow. See [recover-form.test.json](examples/recover-form.test.json) and [recovery rules](docs/design.md#recovery).

The runtime checks browser identity and does not discover personal Chrome. This is browser-session separation, not an operating-system sandbox. Keep CDP on loopback. Trusted fallback scripts run with your local permissions.
Uncertain Save/autosave effects block replay until predeclared UI evidence reconciles them. Missing inputs remain missing. Recovery cannot hide a failed persistence test. Action, model-call, and active-time budgets span all attempts; unknown model spend is reported as unknown.

Jev's MVP lacks support for some frames, shadow roots, canvas, uploads, popup tabs, nested scrolling, and custom keyboard controls. Use [fallback](docs/fallback.md), not a silent pass. Timeouts cannot undo production actions. Cancellation during browser initialization may leave an unregistered tab. `maxSteps` cannot raise upstream limits.
## Agent setup and stored data

The name is not a benchmark claim. CI exercises real Chrome with scripted model decisions. Live-model reliability, production authentication, and performance on your applications remain to be measured. Patchright and visual-agent frameworks are not included.
Keep the three skills: [browser-use](skills/browser-use/SKILL.md), [browser-test](skills/browser-test/SKILL.md), and [setup-fastest-e2e](skills/setup-fastest-e2e/SKILL.md). `fastest-e2e mcp` exposes the same runtime. Arbitrary Python fallback is separately gated by `mcp --allow-scripts`.

## Planned work
Runs, literal goals/inputs, and evidence are local under the configured home. **Do not put secrets in goals or literal input fields.** Use `valueFromEnv` for sensitive input; never allowlist it for recovery. Records expire after 24 hours by default; `fastest-e2e prune` deletes expired data. Screenshots and model context may contain account data. This is not an OS sandbox or complete redaction system.

The [design and implementation checklist](docs/design.md) covers agent ergonomics, scoped/visual fallback, and recovery when live browser state is lost. These are proposals, not current capabilities. A saved run must never imply that a crashed page or half-filled form can be restored.
Reviewed recipe reuse is explicit: `recipe propose`, `recipe approve` with a distinct verified trial, then a task naming that recipe. It reuses UI procedures, not answers or permissions. [Design](docs/design.md) documents limits and quarantine behavior.

## Develop
## Develop and verify

```sh
npm ci
uv sync --project worker --locked
npm run check
uv sync --project worker --locked --python 3.12
uv run --project worker --no-sync python worker/tests/browser_smoke.py
uv run --project worker --no-sync python worker/tests/managed_smoke.py
node test/lifecycle.smoke.mjs
npm run test:browser
```

GitHub Actions use full commit SHAs with version comments. A test catches mutable `uses:` references; Dependabot proposes reviewed Action updates without enabling auto-merge. SHA pinning prevents tag replacement from silently changing the selected action, not every supply-chain risk. See [GitHub's security guidance](https://docs.github.com/en/actions/reference/security/secure-use).

## References
CI uses real Chrome, scripted Jev decisions, and a controlled HTTP model provider with the real Midscene SDK. This proves adapter mechanics, not paid-model quality or compatibility with every production site/client. No performance ranking is claimed. Automatic popup following, arbitrary autonomous file selection, desktop control, and lossless page restoration are not supported.

[rat-stack](https://github.com/joelhooks/rat-stack) inspired shared Effect handlers, exact pins, and typed failures. [Jev Ultrafast](https://github.com/browser-use/jev-ultrafast) supplies the policy and action loop; [Browser Harness](https://github.com/browser-use/browser-harness) supplies CDP and fallback helpers. The human guide follows [show-me](https://github.com/humanlayer/skills/blob/main/plugins/show-me/skills/show-me/SKILL.md). Skill writing follows [unslop](https://github.com/cursor/plugins/blob/main/pstack/skills/unslop/SKILL.md), [writing-for-agents](https://github.com/mattpocock/skills/blob/main/skills/productivity/writing-for-agents/SKILL.md), and Matt Pocock's [setup pattern](https://github.com/mattpocock/skills/blob/main/skills/engineering/setup-matt-pocock-skills/SKILL.md).
Dependencies are locked; Actions use full commit SHAs and read-only normal CI permissions. Pinning prevents tag substitution, not every supply-chain risk. [rat-stack](https://github.com/joelhooks/rat-stack) inspired the shared Effect runtime. [show-me](https://github.com/humanlayer/skills/blob/main/plugins/show-me/skills/show-me/SKILL.md), [unslop](https://github.com/cursor/plugins/blob/main/pstack/skills/unslop/SKILL.md), and [writing-for-agents](https://github.com/mattpocock/skills/blob/main/skills/productivity/writing-for-agents/SKILL.md) inform the guides.

MIT. See [LICENSE](LICENSE).
Loading
Loading