Skip to content

fix(flows): give Garden a 3h budget and start only steps that still fit (cloud#4108) - #132

Merged
khaliqgant merged 4 commits into
mainfrom
rca2/garden-budget-agentrelay.com
Oct 2, 2026
Merged

khaliqgant merged 4 commits into
mainfrom
rca2/garden-budget-agentrelay.com

Conversation

@AgentRelayBot

@AgentRelayBot AgentRelayBot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Fixes the generator half of AgentWorkforce/cloud#4108.

Why

Three Garden runs failed on 2026-10-01, and none failed because the change itself was wrong.

  • bda21b91 used 120.2m of its 2h header budget before the step after its passing checks (step 14). Nothing was published: the kernel refuses every step once the wallclock is spent, including the push.
  • On a repository whose check run takes 11–14 min, one check-and-repair cycle can take an hour. Measured: repairs of 19m, 37m and 44m.
  • Agent steps take no time limit of their own: AgentOptions has no timeout in @relayflows/surface.

What

  • Header wallclock goes from "2h" to "3h" (Cloud's maximum run budget). It comes from FLOW_TIME.headerMinutes, so the header and the plan cannot drift.

  • FLOW_TIME is the time plan, with its arithmetic in a test:

    Step Allowance
    check 15m (its lease)
    repair 45m (≥ the 44m measured)
    review 20m
    fixer 45m
    publish 10m
    setup Cloud spends before the body 5m
  • The generated flow reads the clock through a journaled step (date +%s, so a resumed run replays it). It starts a step only when that step and publishing still fit:

    • a repair needs repair + re-check + base-commit check + publish = 85m left;
    • a review round needs review + publish = 30m;
    • a fix round needs fixer + check + next review + publish = 90m.
  • On a time stop with work done, the work is published, not lost (#4108 ask 2):

    • A repair it cannot afford is skipped, and the existing path opens the PR as a draft with the check report.
    • Out of time after the PR is open: FLOW_TIME_STOP_COMMAND converts it to a draft, comments a note, and the run returns needs_human.

Tests

web/lib/test/flow-budget.test.ts runs the generated flow against a simulated clock and the kernel's budget rule (a step is refused once the charged step time exceeds the header):

  • Red on main (fd5d599 against the old generator): 5 of 6 fail. The bda21b91 replay is refused before the step after its passing checks, the same failure as production.
  • Green (913c9d6), 6/6:
    • bda21b91's measured timings publish;
    • the worst case still publishes, as a draft with a time-stop note, inside the body budget: every check times out, every repair, review and fixer takes its full allowance, the implementer takes 30m;
    • a run with too little time skips the repair and publishes;
    • a fast run is unchanged: one repair, both reviews.
  • flow-onboarding and flow-local assertions are updated from 2h to 3h.
  • I ran the rest of web/lib/test locally on a host at load average ~44. The flow-push-guard and flow-local git tests hit their own 30s timeouts there, so CI is the authority for those.

Limits

  • The allowances are not hard caps: a single agent that runs past its allowance can still exhaust the budget. A real cap needs an agent-step timeout in flows.
  • Already-deployed listeners keep the source they were deployed with, so this only reaches new deployments until a listener is redeployed or its source is patched. Cloud's half is AgentWorkforce/cloud#4113: a deployment's run budget follows the header.

Agent-authored. Per AGENTS.md, this needs two recorded reviews (one from a different agent) and the human/cmo gate. I will not merge it.

🤖 Generated with Claude Code


Note

Medium Risk
Changes generated flow runtime behavior and Cloud run budgets for new deployments; mistakes could skip repairs/reviews or mis-estimate time, though coverage is extensive via simulated budget tests.

Overview
Raises generated Cloud flow wallclock from 2h to 3h (Cloud max), driven by a shared FLOW_TIME plan so the header budget and in-flow guards stay aligned.

The generated flow now tracks remaining time via a journaled date +%s clock (with extra accounting for parallel prototype agents) and skips or short-circuits expensive optional steps—check repairs, base-commit comparison, adversarial reviews, and traditional fix rounds—when their reserved allowance plus publishing would exceed what is left. When reviews cannot finish in time, FLOW_TIME_STOP_COMMAND drafts the PR, comments, and ends with needs_human instead of dying before push.

flow-budget.test.ts simulates kernel budget charging against the generated source (including a replay of run bda21b91); onboarding/local tests expect 3h in the header.

Reviewed by Cursor Bugbot for commit 286cb3e. Bugbot is set up for automated code reviews on this repo. Configure here.


Summary by cubic

Fixes Garden flows dying on their own 2h wallclock budget: once the budget is spent, the kernel refuses every step including the push, so work is lost. Gives the header a 3h budget (Cloud's maximum, from the FLOW_TIME constant) and makes the flow start only steps that still fit the remaining time.

Behavior

  • Repairs, fix rounds, and reviews start only when their allowance plus publishing still fit; the clock is read through a journaled step so a resumed run replays it.
  • Parallel prototypes count as three charges for one wall-clock stretch, matching how the kernel bills them.
  • The base-commit check is skipped when it and publishing no longer fit, and the run reports it was not checked rather than claiming the base also fails.
  • A repair that cannot fit is skipped; the PR still opens as a draft with the check report.
  • Out of time after the PR is open: the flow converts it to a draft, comments a note, and returns needs_human. If no time remains for a fix round, it stops as an unresolved review so the findings go on the PR.

Tests

  • New flow-budget.test.ts replays run bda21b91's measured timings against the kernel's budget rule, including a worst-case run; it is red on the old generator and green here.
  • Fast runs are unchanged: one repair, both review rounds.
  • Onboarding and local starter-kit assertions updated from 2h to 3h.

Written for commit c8db5ce. Summary will update on new commits.

Review in cubic

agentrelaybot added 2 commits October 2, 2026 00:01
Replays run bda21b91's measured step times against the kernel's budget
rule: on the 2h header the run is refused before the step after its
passing checks, as in production (120.2m used before "run-15").

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
…it (cloud#4108)

- Header wallclock 2h -> 3h (Cloud's maximum run budget), from FLOW_TIME.
- The flow reads the clock through a journaled step and starts a repair,
  a fix round or a review only when that step's allowance and publishing
  still fit. Agent steps take no time limit of their own, so this is the
  only cap a flow body can apply.
- Out of time after the pull request is open: push stays, the PR goes
  to draft with a note, and the run asks for a person.
- A repair it cannot afford is skipped; the existing path then opens
  the draft with the check report.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
@coderabbitai

coderabbitai Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 16dca61e-d1f4-4304-b82f-869d4e27f6b9

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployed!

Environment URL
Web https://122a22d8-agentrelay-web.agent-workforce.workers.dev

This is a Cloudflare Workers preview version of this PR's build.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 5 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread web/lib/flow-workflows.ts
Comment thread web/lib/flow-workflows.ts
Comment thread web/lib/test/flow-budget.test.ts
Comment thread web/lib/test/flow-budget.test.ts Outdated
Comment thread web/lib/test/flow-budget.test.ts
Comment thread web/lib/flow-workflows.ts

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 913c9d6. Configure here.

Comment thread web/lib/flow-workflows.ts Outdated
…ew findings on a time stop

Review on #132 (cubic, Cursor Bugbot):
- the budget charges each parallel prototype in full, so the clock adds
  the extra two charges
- the base-commit check is skipped (reported as not checked) when it and
  publishing no longer fit
- no time for a fix round now stops as an unresolved review, so the
  first review's findings go on the pull request
- the test harness gives the body the budget less Cloud's setup, runs
  check-discovery, and models parallel agents on separate clocks

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 2 files (changes from recent commits).

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread web/lib/test/flow-budget.test.ts
Comment thread web/lib/flow-workflows.ts
…prototype charge test

Review round 2 on #132 (cubic):
- a time-skipped base check is "skipped", and the run says there was no
  time to check the base commit instead of claiming it fails too
- the prototype test now fails if parallel time stops counting (it
  asserts no repair starts); each fix was mutation-checked

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants