You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
proposal(ai-sdk): goal-driven step loop with a model-agnostic decide step #2657
PR #2654 showed a real gain and the wrong home for it. The gain: a loop that asks a model one narrow question per step, "which on-screen element advances this goal", over the interactive snapshot, turns N agent turns into one call and cuts decide time per action from seconds to well under one second. The wrong home: a vendor HTTP client, a vendor API key, a pricing constant, and two always-listed CLI commands and MCP tools in core that refuse without the key.
agent-device core is the device side of an agent. Model handling belongs in agent-device/ai-sdk, where the host already brings its own model. That decision is independent of which vendor is behind the head.
The gain is not vendor-specific. On the PR's three sign-in screens plus a real iOS artifact, a small general model (gpt-5.4-nano, reasoning off) with a schema-forced choice among candidate refs chose correctly 16/16 at 620-730 ms median and about 200 input tokens per decision. A decision-only vendor can plug into the same seam from outside this repo.
Proposed shape
Lives under agent-device/ai-sdk next to createAgentDeviceTools, using the optional ai peer. No new CLI command, MCP tool, registry entry, flag or env contract in core.
decide defaults to AI SDK generateObject against the host's model with a schema of { target: enum(refs | none), done, blocked }; a host may pass its own decide(state) => decision;
acts through client.interactions.press / fill with settle: true, every mutation pinned to the snapshot's refsGeneration (ADR 0014);
"did anything change" is read from the settle diff and tail on SettleObservation, not from a second snapshot;
text is supplied, never generated; sensitive values go through the ADR 0017 channel (recordAs, AD_VAR_*), not a new --input or env family;
returns the step table (screen, decision, outcome, snapshot/decide/action ms) and a typed status: done, blocked, escalated, max-steps.
Out of scope: scroll, back, alert and gesture planning; multi-screen plans; text generation.
Completion conditions
examples/sdk/goal-loop.ts or the ai-sdk export drives a sign-in flow on an iOS simulator and on an Android emulator with a host-configured model, with the step table and the --debug request log showing ~sN pinned refs attached.
Unit coverage over a fake device port and a fake decide, with fixtures that carry the errors production emits.
Docs: one section in website/docs/docs/ai-sdk.md.
Open decisions
In-tree ai-sdk export, or examples/sdk/goal-loop.ts first.
Confidence floor and unproductive-step limit defaults; whether blocked is ever terminal on its own.
Whether kind-based candidate filtering is enough or the loop needs a hittable / interactionBlocked gate.
Dependencies
Blocked by: #2656. Related: #2634 (fill on fields that normalize their input), ADR 0014, ADR 0017.
Purpose
PR #2654 showed a real gain and the wrong home for it. The gain: a loop that asks a model one narrow question per step, "which on-screen element advances this goal", over the interactive snapshot, turns N agent turns into one call and cuts decide time per action from seconds to well under one second. The wrong home: a vendor HTTP client, a vendor API key, a pricing constant, and two always-listed CLI commands and MCP tools in core that refuse without the key.
agent-device core is the device side of an agent. Model handling belongs in
agent-device/ai-sdk, where the host already brings its own model. That decision is independent of which vendor is behind the head.The gain is not vendor-specific. On the PR's three sign-in screens plus a real iOS artifact, a small general model (
gpt-5.4-nano, reasoning off) with a schema-forced choice among candidate refs chose correctly 16/16 at 620-730 ms median and about 200 input tokens per decision. A decision-only vendor can plug into the same seam from outside this repo.Proposed shape
agent-device/ai-sdknext tocreateAgentDeviceTools, using the optionalaipeer. No new CLI command, MCP tool, registry entry, flag or env contract in core.runGoal({ client, goal, model | decide, inputs, maxSteps, minConfidence, onStep }):kind, name, value, enabled, ref (needs feat(snapshot): carry the presenter's platform-neutral role on structured snapshot nodes #2656);decidedefaults to AI SDKgenerateObjectagainst the host's model with a schema of{ target: enum(refs | none), done, blocked }; a host may pass its owndecide(state) => decision;client.interactions.press/fillwithsettle: true, every mutation pinned to the snapshot'srefsGeneration(ADR 0014);diffandtailonSettleObservation, not from a second snapshot;recordAs,AD_VAR_*), not a new--inputor env family;done,blocked,escalated,max-steps.Completion conditions
examples/sdk/goal-loop.tsor theai-sdkexport drives a sign-in flow on an iOS simulator and on an Android emulator with a host-configured model, with the step table and the--debugrequest log showing~sNpinned refs attached.decide, with fixtures that carry the errors production emits.website/docs/docs/ai-sdk.md.Open decisions
ai-sdkexport, orexamples/sdk/goal-loop.tsfirst.blockedis ever terminal on its own.kind-based candidate filtering is enough or the loop needs ahittable/interactionBlockedgate.Dependencies
Blocked by: #2656. Related: #2634 (fill on fields that normalize their input), ADR 0014, ADR 0017.