Repository navigation
fix(flows): give Garden a 3h budget and start only steps that still fit (cloud#4108) - #132
Conversation
Replays run bda21b91's measured step times against the kernel's budget rule: on the 2h header the run is refused before the step after its passing checks, as in production (120.2m used before "run-15"). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
…it (cloud#4108) - Header wallclock 2h -> 3h (Cloud's maximum run budget), from FLOW_TIME. - The flow reads the clock through a journaled step and starts a repair, a fix round or a review only when that step's allowance and publishing still fit. Agent steps take no time limit of their own, so this is the only cap a flow body can apply. - Out of time after the pull request is open: push stays, the PR goes to draft with a note, and the run asks for a person. - A repair it cannot afford is skipped; the existing path then opens the draft with the check report. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Preview deployed!
This is a Cloudflare Workers preview version of this PR's build. |
There was a problem hiding this comment.
All reported issues were addressed across 5 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 913c9d6. Configure here.
…ew findings on a time stop Review on #132 (cubic, Cursor Bugbot): - the budget charges each parallel prototype in full, so the clock adds the extra two charges - the base-commit check is skipped (reported as not checked) when it and publishing no longer fit - no time for a fix round now stops as an unresolved review, so the first review's findings go on the pull request - the test harness gives the body the budget less Cloud's setup, runs check-discovery, and models parallel agents on separate clocks Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff
There was a problem hiding this comment.
All reported issues were addressed across 2 files (changes from recent commits).
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
…prototype charge test Review round 2 on #132 (cubic): - a time-skipped base check is "skipped", and the run says there was no time to check the base commit instead of claiming it fails too - the prototype test now fails if parallel time stops counting (it asserts no repair starts); each fix was mutation-checked Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Session-Id: 3b31cc74-0c43-4890-9650-f2be3566b7ff

Fixes the generator half of AgentWorkforce/cloud#4108.
Why
Three Garden runs failed on 2026-10-01, and none failed because the change itself was wrong.
AgentOptionshas no timeout in@relayflows/surface.What
Header
wallclockgoes from "2h" to "3h" (Cloud's maximum run budget). It comes fromFLOW_TIME.headerMinutes, so the header and the plan cannot drift.FLOW_TIMEis the time plan, with its arithmetic in a test:The generated flow reads the clock through a journaled step (
date +%s, so a resumed run replays it). It starts a step only when that step and publishing still fit:On a time stop with work done, the work is published, not lost (#4108 ask 2):
FLOW_TIME_STOP_COMMANDconverts it to a draft, comments a note, and the run returnsneeds_human.Tests
web/lib/test/flow-budget.test.tsruns the generated flow against a simulated clock and the kernel's budget rule (a step is refused once the charged step time exceeds the header):flow-onboardingandflow-localassertions are updated from 2h to 3h.web/lib/testlocally on a host at load average ~44. Theflow-push-guardandflow-localgit tests hit their own 30s timeouts there, so CI is the authority for those.Limits
Agent-authored. Per AGENTS.md, this needs two recorded reviews (one from a different agent) and the human/cmo gate. I will not merge it.
🤖 Generated with Claude Code
Note
Medium Risk
Changes generated flow runtime behavior and Cloud run budgets for new deployments; mistakes could skip repairs/reviews or mis-estimate time, though coverage is extensive via simulated budget tests.
Overview
Raises generated Cloud flow wallclock from 2h to 3h (Cloud max), driven by a shared
FLOW_TIMEplan so the header budget and in-flow guards stay aligned.The generated flow now tracks remaining time via a journaled
date +%sclock (with extra accounting for parallel prototype agents) and skips or short-circuits expensive optional steps—check repairs, base-commit comparison, adversarial reviews, and traditional fix rounds—when their reserved allowance plus publishing would exceed what is left. When reviews cannot finish in time,FLOW_TIME_STOP_COMMANDdrafts the PR, comments, and ends withneeds_humaninstead of dying before push.flow-budget.test.tssimulates kernel budget charging against the generated source (including a replay of run bda21b91); onboarding/local tests expect3hin the header.Reviewed by Cursor Bugbot for commit 286cb3e. Bugbot is set up for automated code reviews on this repo. Configure here.
Summary by cubic
Fixes Garden flows dying on their own 2h wallclock budget: once the budget is spent, the kernel refuses every step including the push, so work is lost. Gives the header a 3h budget (Cloud's maximum, from the
FLOW_TIMEconstant) and makes the flow start only steps that still fit the remaining time.Behavior
needs_human. If no time remains for a fix round, it stops as an unresolved review so the findings go on the PR.Tests
flow-budget.test.tsreplays run bda21b91's measured timings against the kernel's budget rule, including a worst-case run; it is red on the old generator and green here.Written for commit c8db5ce. Summary will update on new commits.