test(e2e): split model and world suites - #1387
Conversation
The e2e matrix crossed every fixture with every model in both workflows
(84 jobs/PR), so full runs rarely passed. Split the coverage:
- e2e-local is the model suite: real matrix models against the local
world; only model-sensitive fixtures ("e2e": { "modelMatrix":
"full" }) run on every model, the rest run once on the default.
- e2e-vercel is the world suite: every fixture deploys once with
deterministic mock models (EVE_E2E_MODEL=mock) and excludes
real-model-tagged evals, proving world infrastructure without
live-model flake. Additional worlds (e.g. Postgres) follow the same
shape.
Adds the private @eve-e2e/config harness package (shared model, world,
and judge resolution for fixtures), a repeatable --exclude-tag flag for
eve eval (exclusion that removes every match exits 0), a two-output
fixture discovery script, and a scripted mock responder for
agent-tools-sandbox so the redeploy eval keeps coverage. Evals not yet
verified under mocks carry the real-model tag; untagging them per
fixture is the migration path for world-suite coverage.
Signed-off-by: Andrew Barba <barba@hey.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Bundle + Package Summary:
|
| Area | Metric | Baseline | Current | Delta |
|---|---|---|---|---|
| Package | Packed tarball | 7.62 MB | 7.62 MB | +1.0 kB |
| Package | Unpacked publish size | 28.80 MB | 28.80 MB | +1.8 kB |
| Package | Installed footprint | 91.28 MB | 91.28 MB | +1.8 kB |
| Package | Published files | 2876 | 2878 | +2 |
| Package | Installed files | 6665 | 6667 | +2 |
| Runtime | Unique function payloads | 2 | 2 | 0 |
| Runtime | Total function bytes | 16.81 MB | 16.81 MB | -40 B ✅ |
| Runtime | Public routes | 11 | 11 | 0 |
Changed function payloads vs main (255027e) (2)
| Function | Status | Baseline | Current | Delta | Route changes |
|---|---|---|---|---|---|
functions/__server.func |
changed | 8.41 MB | 8.41 MB | -20 B ✅ | none |
functions/.well-known/workflow/v1/flow.func |
changed | 8.41 MB | 8.41 MB | -20 B ✅ | none |
eve init install
| Metric | Baseline | Current | Delta |
|---|---|---|---|
| Installed footprint | 129.68 MB | 129.68 MB | +1.8 kB |
| Installed packages | 128 | 128 | 0 |
| dependencies | 4 | 4 | 0 |
| devDependencies | 2 | 2 | 0 |
| Dependency package bytes | 43.14 MB | 43.14 MB | +1.8 kB |
| devDependency package bytes | 5.04 MB | 5.04 MB | 0 B ➖ |
Build Metadata
- Preset:
vercel - Nitro:
nitro@3.0.260610-beta - Output directory:
apps/fixtures/weather-agent/.vercel/output - Build metadata timestamp: 2026-07-30T18:48:01.764Z
- Route aliases: 11 public, 1 internal (12 total aliases)
- Vercel routes in config: 14
- Severity legend: 🔴 dominant/large, 🟠 notable, 🟡 watch, ⚪ small
Package Drill-Down
Package Details
- Package:
eve@0.28.0 - Package directory:
packages/eve - Tarball: 7.62 MB (
eve-0.28.0.tgz) - Unpacked payload: 28.80 MB across 2878 published files
- Installed footprint: 91.28 MB across 6667 installed files
- Installed root package: 27.45 MB
- Installed dependencies: 63.83 MB
- Runtime dependencies: 2
- Peer dependencies: 5 (4 optional)
Installed footprint is measured from an isolated temporary npm install of the packed tarball.
Heavy installed dependencies
eve: 27.45 MB (30.1%)@rolldown/binding-linux-x64-gnu: 19.28 MB (21.1%)@rolldown/binding-wasm32-wasi: 10.66 MB (11.7%)@napi-rs/wasm-runtime: 6.56 MB (7.2%)ai: 6.53 MB (7.2%)
Publish payload breakdown
Published file size
🔴 dist/src/compiled/shadcn-registry/index.js [########################] 13.15 MB 45.7%
🟠 dist/src/compiled/experimental-ai-sdk-code-mo... [###.....................] 1.51 MB 5.2%
🟡 dist/src/compiled/@vercel/sandbox/index.js [#.......................] 632.5 kB 2.2%
🟡 dist/src/compiled/_chunks/workflow/undici-DWL... [#.......................] 502.4 kB 1.7%
🟡 dist/src/compiled/@chat-adapter/slack/index.js [#.......................] 440.5 kB 1.5%
🔴 Other published files [#######################.] 12.57 MB 43.6%
Installed footprint breakdown
Installed package size
🔴 eve [########################] 27.45 MB 30.1%
🔴 @rolldown/binding-linux-x64-gnu [#################.......] 19.28 MB 21.1%
🔴 @rolldown/binding-wasm32-wasi [#########...............] 10.66 MB 11.7%
🔴 @napi-rs/wasm-runtime [######..................] 6.56 MB 7.2%
🔴 ai [######..................] 6.53 MB 7.2%
🔴 zod [####....................] 5.07 MB 5.6%
🔴 Other installed packages [##############..........] 15.73 MB 17.2%
Runtime dependencies (2)
| Package | Range | Notes |
|---|---|---|
nitro |
3.0.260610-beta |
|
undici |
8.9.0 |
Peer dependencies (5)
| Package | Range | Notes |
|---|---|---|
@opentelemetry/api |
^1.0.0 |
optional peer |
ai |
catalog: |
|
braintrust |
^3.0.0 |
optional peer |
just-bash |
^3.0.0 |
optional peer |
microsandbox |
^0.5.0 |
optional peer |
eve init install drill-down
eve init install details
- Command:
eve init my-agent - Package manager:
npm - Installed footprint: 129.68 MB across 8535 installed files
- Installed packages: 128 total (122 transitive-only)
- dependencies: 4 direct packages totaling 43.14 MB
- devDependencies: 2 direct packages totaling 5.04 MB
- Other transitive package files: 81.50 MB
Installed footprint is measured from an isolated temporary eve init my-agent using the current packed eve tarball.
Heavy installed dependencies
@typescript/typescript-linux-x64: 27.95 MB (21.5%)eve: 27.45 MB (21.2%)@rolldown/binding-linux-x64-gnu: 19.28 MB (14.9%)@rolldown/binding-wasm32-wasi: 10.66 MB (8.2%)zod: 9.02 MB (7.0%)
Installed footprint breakdown
Installed package size
🔴 @typescript/typescript-linux-x64 [########################] 27.95 MB 21.5%
🔴 eve [########################] 27.45 MB 21.2%
🔴 @rolldown/binding-linux-x64-gnu [#################.......] 19.28 MB 14.9%
🔴 @rolldown/binding-wasm32-wasi [#########...............] 10.66 MB 8.2%
🔴 zod [########................] 9.02 MB 7.0%
🔴 @napi-rs/wasm-runtime [######..................] 6.56 MB 5.1%
🔴 ai [######..................] 6.53 MB 5.0%
🔴 Other installed packages [###################.....] 22.23 MB 17.1%
dependencies (4)
| Package | Range | Installed size | Share |
|---|---|---|---|
@vercel/connect |
0.4.2 |
135.8 kB | 0.1% |
ai |
^7.0.38 |
6.53 MB | 5.0% |
eve |
file:eve-0.28.0.tgz |
27.45 MB | 21.2% |
zod |
4.4.3 |
9.02 MB | 7.0% |
devDependencies (2)
| Package | Range | Installed size | Share |
|---|---|---|---|
@types/node |
24.x |
2.54 MB | 2.0% |
typescript |
7.0.2 |
2.50 MB | 1.9% |
Function Drill-Down
Payload Size Graph
Unique function payload size and share of total
🔴 functions/.well-known/workflow/v1/flow.func [########################] 8.41 MB 50.0%
🔴 functions/__server.func [########################] 8.41 MB 50.0%
Top Function Payloads
🟠 functions/.well-known/workflow/v1/flow.func • 1 public route • 8.41 MB
| Metric | Value |
|---|---|
| Public routes | /.well-known/workflow/v1/flow |
| Runtime | nodejs24.x |
| Handler | index.mjs |
| Payload | 8.41 MB |
| Function files | 8.41 MB across 43 files |
| Traced dependencies | 0 B |
| Signal | 🟠 Bundled file index.mjs is 2.31 MB (27.5%) |
🟠 🔎 Dependency Analysis
📦 Bundled files:
Bundled file size
🟠 index.mjs [########################] 2.31 MB 27.5%
🟠 _chunks/runtime-artifacts.mjs [################........] 1.59 MB 19.0%
🟡 _libs/undici.mjs [##########..............] 980.5 kB 11.7%
🟡 _chunks/sandbox.mjs [########................] 768.8 kB 9.1%
🟡 _libs/@ai-sdk/gateway+[...].mjs [####....................] 432.8 kB 5.1%
🟠 Other bundled files [########################] 2.32 MB 27.6%
🧾 Vercel Config
{
"handler": "index.mjs",
"launcherType": "Nodejs",
"shouldAddHelpers": false,
"supportsResponseStreaming": true,
"runtime": "nodejs24.x",
"maxDuration": "max",
"experimentalTriggers": [
{
"type": "queue/v2beta",
"topic": "__eve776561746865722d6167656e74_wkf_workflow_*",
"consumer": "default",
"retryAfterSeconds": 5,
"initialDelaySeconds": 0
}
],
"environment": {
"WORKFLOW_PRECONDITION_GUARD": "1"
}
}🟠 functions/__server.func • 10 public routes, 1 internal alias • 8.41 MB
| Metric | Value |
|---|---|
| Public routes | //eve/v1/callback/[token]/eve/v1/connections/[name]/callback/[token]/eve/v1/health/eve/v1/info/eve/v1/session/eve/v1/session/[sessionId]/eve/v1/session/[sessionId]/cancel/eve/v1/session/[sessionId]/stream/eve/v1/session/reset |
| Internal aliases | /__server |
| Runtime | nodejs24.x |
| Handler | index.mjs |
| Payload | 8.41 MB |
| Function files | 8.41 MB across 43 files |
| Traced dependencies | 0 B |
| Signal | 🟠 Bundled file index.mjs is 2.31 MB (27.5%) |
🟠 🔎 Dependency Analysis
📦 Bundled files:
Bundled file size
🟠 index.mjs [########################] 2.31 MB 27.5%
🟠 _chunks/runtime-artifacts.mjs [################........] 1.59 MB 19.0%
🟡 _libs/undici.mjs [##########..............] 980.5 kB 11.7%
🟡 _chunks/sandbox.mjs [########................] 768.8 kB 9.1%
🟡 _libs/@ai-sdk/gateway+[...].mjs [####....................] 432.8 kB 5.1%
🟠 Other bundled files [########################] 2.32 MB 27.6%
🧾 Vercel Config
{
"handler": "index.mjs",
"launcherType": "Nodejs",
"shouldAddHelpers": false,
"supportsResponseStreaming": true,
"runtime": "nodejs24.x"
}Build Timing: e2e/fixtures/agent-tools-sandbox
This is an informational timing measurement inside eve build, from preflight through publication. Output-size measurement and profile writing are excluded.
Build mode: deployable Vercel build with sandbox template prewarm included.
- Build pipeline: 2.04 s -> 1.95 s (-87.6 ms) vs
main (255027e). - Timing is informational: shared GitHub runners are too variable for a hard timing budget.
Detailed phase timings vs `main (255027e)`
| Phase | Baseline | Current | Delta |
|---|---|---|---|
extension.check |
6.0 ms | 9.2 ms | +3.2 ms |
project.resolve |
3.5 ms | 1.5 ms | -2.0 ms |
workspace.create |
1.1 ms | 1.3 ms | +0.2 ms |
host.prepare |
283.8 ms | 213.9 ms | -69.9 ms |
vercel.service-prefix.resolve |
2.0 ms | 2.2 ms | +0.2 ms |
nitro.create |
195.0 ms | 193.4 ms | -1.6 ms |
sandbox.prewarm |
233.4 ms | 214.3 ms | -19.1 ms |
nitro.cache.prepare |
0.2 ms | 0.2 ms | 0.0 ms |
nitro.prepare |
0.7 ms | 0.7 ms | 0.0 ms |
nitro.public-assets |
0.7 ms | 0.7 ms | 0.0 ms |
nitro.prerender |
0.5 ms | 0.5 ms | 0.0 ms |
nitro.bundle |
1.28 s | 1.29 s | +1.4 ms |
nitro.cache.write |
0.4 ms | 0.4 ms | 0.0 ms |
vercel.workflow-function.materialize |
20.1 ms | 20.1 ms | 0.0 ms |
agent-summary.emit |
0.5 ms | 0.5 ms | 0.0 ms |
nitro.close |
0.2 ms | 0.2 ms | 0.0 ms |
output.publish |
3.1 ms | 3.1 ms | 0.0 ms |
workspace.remove |
1.8 ms | 1.7 ms | -0.1 ms |
The main ruleset already requires an e2e-postgres aggregate check, but no workflow reported it, leaving every PR blocked on a permanently expected status. Add the Postgres world suite following the same shape as e2e-vercel: one leg per fixture from world_matrix, deterministic mock models (EVE_E2E_MODEL=mock), and real-model-tagged evals excluded. Each leg starts a PostgreSQL service container, bootstraps the @workflow/world-postgres schema, builds the fixture with EVE_E2E_WORKFLOW_WORLD=@workflow/world-postgres, runs mock-compatible evals against a local production server, and asserts the traffic produced Postgres-backed workflow runs. Fixtures carry @workflow/world-postgres so the world module resolves at build time. Verified locally against postgres:18-alpine: agent-workflow-stress and agent-basic-runtime pass with hundreds of workflow runs persisted. Signed-off-by: Andrew Barba <barba@hey.com>
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
Salvaged from #1367: pin the image-attachment reply so the transcript projection assertion carries the coverage (the eval now runs in the world suites and passes under EVE_E2E_MODEL=mock), phrase the sleeper delegation as an explicit directive a scripted responder can drive, and serialize agent-subagents evals so deterministic turns do not compete with the child workflows they test. Signed-off-by: Andrew Barba <barba@hey.com>
Signed-off-by: Andrew Barba <barba@hey.com> # Conflicts: # pnpm-lock.yaml
List mode previously sat below the all-excluded early return, so `--list --json --exclude-tag <tag>` printed a prose notice instead of JSON when exclusion removed every eval. A list command should list the selected (possibly empty) set: suite runners can now probe whether anything would run with `--list --json`. Signed-off-by: Andrew Barba <barba@hey.com>
Every world-suite leg asserted at least one Postgres-backed workflow run, but 15 of 20 fixtures currently execute nothing there: either all evals carry the real-model tag (excluded, no junit written) or the only untagged eval self-skips at runtime (agent-tools-sandbox's redeploy responder without Vercel credentials). Read executed = tests - skipped from the junit report and skip the assertion when nothing ran; those legs still prove the fixture builds, boots, and reports healthy against the Postgres world. Signed-off-by: Andrew Barba <barba@hey.com>
The stamped-id contract (#protocol/event-id.js) promises mint-order sorting within a process, not across the separate steps of one session: a turn's events are appended by different steps whose durable-stream writes can interleave behind minting. The rewind test asserted globally sorted ids anyway and flaked on windows-latest, where slower writes widen that window. Uniqueness, shape, and rewind stability remain asserted. Signed-off-by: Andrew Barba <barba@hey.com>
The matrix models lived inside a Node heredoc in a bash discovery script — an unfindable place to edit, and worlds were not registered anywhere at all. e2e/matrix.json is now the single registry: models drive the model suite as before, and each registered world backs a world_matrix_<name> output whose optional package reaches the world workflow as matrix.world_package (EVE_E2E_WORKFLOW_WORLD). The script becomes a plain .mjs that validates the registry (names are check identifiers) and emits the matrices. Adding a model is one registry line; adding a world is one registry line plus a workflow that copies the postgres shape. Signed-off-by: Andrew Barba <barba@hey.com>
OwenKephart
left a comment
There was a problem hiding this comment.
naming's a bit odd for --tag / --exclude-tag, my guess is that in the future we'll want to have more generic --include / --exclude that can work based off of tags / paths / etc. but seems fine for now 🫡
Why
The e2e matrix crossed every fixture with every model in both workflows (84 jobs/PR), so a full run rarely passed and every flake blocked merges. Per [the Slack discussion], most changes cannot plausibly break one flow on one model deterministically, and only a handful of fixtures vary per provider — so the cross product mostly bought flake, not coverage.
What
Splits e2e into two suites:
e2e-local.yml"e2e": { "modelMatrix": "full" }) run on both, the rest run once on the defaulte2e-vercel.ymlmockModel()viaEVE_E2E_MODEL=mock;real-model-tagged evals excludede2e-postgres.yml@workflow/world-postgres(service container)84 flaky jobs/PR → 64 jobs of which 40 are deterministic. Aggregate check names
e2e-localande2e-vercelare unchanged; the newe2e-postgresaggregate satisfies the check the main ruleset already required (previously nothing reported it, leaving every PR blocked on a permanently "expected" status). Any additional world follows the same shape: consumeworld_matrix, setEVE_E2E_MODEL=mock+EVE_E2E_WORKFLOW_WORLD, run with--exclude-tag real-model.The Postgres suite was validated locally against
postgres:18-alpine: bootstrap → build →eve start→ mock evals green, with hundreds of workflow runs persisted (workflow.workflow_runs).Pieces
@eve-e2e/config(e2e/fixtures/e2e-config, extracted from chore(e2e): coverage - split model and workflow world suites #1367's design so that PR rebases down to its postgres-specific parts):e2eAgentConfig/e2eSubagentConfig/e2eModel/e2eJudgeModelresolve the matrix model, the mock sentinel, and theEVE_E2E_WORKFLOW_WORLDoverride in one place. Adopted across all 20 fixtures.eve eval --exclude-tag <tag...>: repeatable exclusion applied after--tag; a run where exclusion removes every match exits 0 ("this suite does not apply here"), while a no-match--tagstays exit 2. Unit-tested, documented, changeset included.model_matrix(fixture × model legs) andworld_matrix(one leg per fixture); the model registry lives in the script.real-modeltag: evals whose assertions need a live model (judges, cache metrics, free-form tool planning). 90 evals start tagged; 10 already run mock-clean (verified locally withEVE_E2E_MODEL=mock): agent-workflow-stress, agent-session-timeout, agent-channels, agent-model (2/3), agent-basic-runtime (3/6), and the sandbox redeploy eval via a scripted directive-driven responder (it has no model-suite fallback). Untagging per fixture is the migration path for world-suite coverage.e2e/README.mdrewritten around the two suites; stale AGENTS.md fixture reference fixed; leftoveragent-compaction-regressions-*shell dirs removed.Testing
pnpm fmt/pnpm lint/pnpm guard:invariants/pnpm docs:checkEVE_E2E_MODEL=mock pnpm exec eve eval --strict --exclude-tag real-model) green for the untagged fixturesagent-tools-sandboxredeploy responder is the only untagged eval that needs Vercel credentials to validate — if it misbehaves, tag itreal-modeland iterate.