Answer ambiguous requests with choices and sourced comparison cards - #42
Conversation
|
I like that the server has a real durable head ( I think the mobile side may currently have a second source of truth, though.
So this sequence looks possible:
At that point the UI says “Earlier choices” and disables P, while Could we add a regression for Either the durable head should be expired when that new turn is admitted, or the UI should project panel availability from durable Jev state rather than independently reconstructing it from transcript order. That seems especially important for Rich Threads / cross-device replay: ephemeral controls should be a projection of the durable interaction state, not a competing state machine. AI-use note: I used an AI assistant to trace the Jev head lifecycle and mobile transcript projection and draft this review; I verified the described paths against the current PR head before posting. |
|
One integration-boundary question for live mode, especially now that this head includes a real live-Jev recording: could we document exactly what leaves OpenMuse when
That The docs are now clear about which recordings use real Jev and that Could we add a short live-mode privacy/config note? Something along the lines of: enabling live Jev sends the request/context and prepared candidate data to the configured Jev provider for control selection/scoring; source data remains evidence and Jev cannot approve or execute actions. I think that would make the otherwise nice authority boundary here explicit to operators. AI-use note: I used an AI assistant to inspect the live adapter payload and current documentation and draft this review; I verified the listed fields against the current PR head before posting. |
Review: Jev usageChecked against the Verdict: the SDK calls are correct and the integration keeps code in control. The gaps are in how the answers are used. What's right
Should fix
Worth tightening
Also: the description says "Jev chooses the prepared control", but the agent chooses the control and Jev decides whether to show it. Suggested wording: "Jev decides whether to show the agent's prepared choices and ranks them." I'll push fixes for 1–9 to this branch. |
- Only let Jev override the agent's prepared control when its choice confidence is at least 0.5; an uncertain answer keeps the agent's control. - Rank options by rounded rubric level with stable ties, since fractional scores between levels are not calibrated finely enough to order. - Send the person's latest message as `userMessage` in state, alongside the agent's summary, so Jev judges independently of the agent's framing. - Refer to options by position (`options[i]`) in question text and keep page-derived labels in state only. - Log provider failures with status and request ID, and tell the agent not to retry non-retryable 4xx responses such as 401. - Build the adapter once per server so live mode shares one TypeSafe client, with a 5s timeout and one retry instead of the SDK's ~30s worst case. - Keep the default Jev model in one place (config.ts). - Test against SDK-shaped answers and add contract tests that drive the real TypeSafeClient through a fake fetch (request shape, retry, 401, abort).
|
Pushed cec18d7, which addresses 1–9 from the review above.
Checks: Not re-run: the live TypeSafe recording. Ranking now rounds to rubric levels and the prompt wording changed, so the Rocky Shore result in the recording should be re-checked with a real key. The sample adapter still never returns |
The mobile transcript marks a panel stale as soon as a later user message
appears, but the server only expired the head on RUN_FINISHED. A follow-up
that failed or was cancelled left the UI showing "Earlier choices" while
select() still accepted the panel.
- Expire the current panel when a new user turn is admitted, before the
agent runs, so the durable head matches the transcript on every device.
Runs resuming after a tool result are not new turns and keep the panel.
- Record which turn retired the panel so that turn can still refine it
("Something hands-on"); a later turn cannot, and selection stays rejected.
- Regression tests: panel -> ordinary turn -> failed run, and -> cancelled
run, both reject the old selection; resumed runs keep the panel; only
the retiring turn may refine.
- Document exactly what JEV_MODE=live sends to TypeSafe and what Jev can
and cannot do, with pointers from SECURITY.md and .env.example.
|
@kvnloo thanks, both addressed in 97dfa78. Two sources of truth. The server is now authoritative, and it agrees with the transcript. The current panel is retired when a new user turn is admitted, before the agent runs, instead of on
Regressions in I kept the transcript projection on mobile rather than adding a durable-state endpoint. With the server matching it, both stay consistent across devices and replay. Live-mode disclosure. New section What live mode sends to TypeSafe. It lists each field ( Locally: typecheck, lint, server build, and 225/225 tests pass. |
Resolve package.json and pnpm-lock.yaml: keep main's TanStack AI dependencies (CopilotKit#46) and add @typesafe-ai/sdk.
Resolve tests/config.test.ts: keep both the Jev mode test and main's shadowed-.env test (CopilotKit#57).
上游 5 个提交:JEV(CopilotKit#42)、无 scheme worker 地址(CopilotKit#92)、web 输入框焦点(CopilotKit#89)、 worker 测试健壮性(CopilotKit#90)、render 蓝图(CopilotKit#86)。 冲突解决(4 文件 9 块)与本地适配: - conversation.ts:以上游 JEV 结构(runInternal/choices 分支/noteEvidence)为基座, 换回我们的中性工具层(forAgUi + chatTools),并把 noteEvidence 改成经 ChatToolContext 回调; prompt 用我们的 chatInstructions() + JEV 段 + computerInstructions;恢复 maxSteps=10 与 sample 分支的 threadId 注入 - agent.ts:保留我们的 mastra 引擎分支,两处 ConversationAgent 注入 sharedJevAdapter() - chat.tsx:我们的卡片(搜索/浏览器操作/审批)与 JEV 卡片并存,去重 delegate_task 渲染器 - config.ts:去掉 cherry-pick 造成的 browserWorkerUrl 重复定义 - mastra-engine.ts:JEV 工具与指令补到 mastra 路径(上游只挂自带引擎,我们线上跑 mastra) - jev/tools.ts:抽出 presentChoicesSpec 中性规格,两条引擎共用同一份校验逻辑 验证:pnpm test 321/321 通过(含上游 JEV 全部测试)
jev-live-web.mp4
Live Jev recording, 83 seconds. The cards are labeled Live Jev · model decisions. AI Mock scripts the agent's steps so the fictional school-trip scenario repeats; the real browser worker reads the aquarium's public pages. After "Something hands-on", Jev's ranking moves Rocky Shore to first place. A scripted-decision sample runs without a TypeSafe key. Reproduction and provenance.
The problem
When a request has several reasonable next steps, OpenMuse can only answer in prose. Someone preparing for a school trip gets a paragraph listing "complete the permission slip, review the details, or explore exhibits" and has to type one back. When they ask the agent to compare options, the comparison is also prose, with nothing tying each claim to the page it came from, and a follow-up like "something hands-on" means the agent writes the whole comparison again.
So people either accept an unsourced summary or open every page themselves to check it. Neither survives a second device or a reload, because nothing records which options were offered or which one was chosen.
The approach
The agent gets one new tool,
present_choices. It prepares either clarification buttons or up to three sourced comparison cards, and a click comes back as a structured choice the agent continues from. The agent owns research and candidate copy. TypeSafe Jev makes two narrow decisions: whether to show the agent's prepared cards or answer in prose, and how well each candidate fits the user's latest message. Both go in one batched call. Code keeps control: Jev can only pick from the agent's control or prose, it overrides the agent only when its confidence is at least 0.5, and options are ranked by whole rubric level with ties kept in the agent's order, because fractional scores between levels are not calibrated finely enough to order.Evidence before comparison. In live mode a comparison is refused unless every source URL was read successfully in the same turn, and each label, source title and detail appears in that page's text, ignoring case and spacing. A mail-based choice requires the thread to have been read. This costs the agent a browse per source, which is the point: a card cannot cite a page nobody opened.
The server owns panel state. Panels and selections are versioned records under
jev_threads, updated by compare-and-swap, so stale, cross-thread and conflicting clicks are rejected and a failed selection can be retried. A new user turn retires the current panel before the agent runs. That matches the chat transcript, where any later user message makes earlier choices stale, even if the turn then fails or is cancelled. The retiring turn may still refine the panel, which is how "something hands-on" re-ranks the stored candidates without new browsing. A run that resumes after a tool result is not a new turn and keeps the panel.Jev sees the person, not the agent's paraphrase. State carries the user's own latest message next to the agent's summary, and questions refer to options by position so page-derived labels stay out of the instructions.
Bounded provider cost. One
TypeSafeClientis shared per server, with a 5-second timeout and one retry: about 10 seconds worst case instead of the SDK's default of 30 or more. Failures log only HTTP status and request ID. A 401 or other 4xx that cannot succeed tells the agent not to retry and to answer in prose.JEV_MODEdefaults tooff.sampleuses a scripted scorer and needs no key.liverequires a server-onlyTYPESAFE_API_KEYand refuses to start without one. The same card component renders in web and native chat.What is not covered
Verification
Locally on the merged head: root, mobile and server TypeScript checks, Biome, the server build, and 275 of 275 tests. CI passes all seven jobs: quality, the web, iOS and Android builds, browser, browser-container and computer-container.
tests/jev.test.ts: the adapter against SDK-shaped answers, including the confidence gate, rubric-level ranking, positional references and failure logging. Contract tests drive the realTypeSafeClientthrough a fakefetch: request shape, one retry on 503, no retry on 401, abort.tests/jev-persistence.test.ts: versioned panels, compare-and-swap selection, stale and cross-thread rejection, refinement and retry.tests/conversation-jev.test.ts: evidence checks for mail, redirects, invented labels, titles and details. Also a panel followed by a failed or cancelled turn rejecting the old selection, a resumed run keeping it, and only the retiring turn refining it.apps/mobile/test/jev-actions.test.ts: the transcript projection that decides which card accepts input.tests/config.test.ts,tests/demo-model.test.ts: mode validation and the scripted demo.