Skip to content

Answer ambiguous requests with choices and sourced comparison cards - #42

Merged
jerelvelarde merged 22 commits into
CopilotKit:mainfrom
jerelvelarde:jerel/jev-generative-ui
Sep 29, 2026
Merged

jerelvelarde merged 22 commits into
CopilotKit:mainfrom
jerelvelarde:jerel/jev-generative-ui

Conversation

@jerelvelarde

@jerelvelarde jerelvelarde commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator
jev-live-web.mp4

Live Jev recording, 83 seconds. The cards are labeled Live Jev · model decisions. AI Mock scripts the agent's steps so the fictional school-trip scenario repeats; the real browser worker reads the aquarium's public pages. After "Something hands-on", Jev's ranking moves Rocky Shore to first place. A scripted-decision sample runs without a TypeSafe key. Reproduction and provenance.

The problem

When a request has several reasonable next steps, OpenMuse can only answer in prose. Someone preparing for a school trip gets a paragraph listing "complete the permission slip, review the details, or explore exhibits" and has to type one back. When they ask the agent to compare options, the comparison is also prose, with nothing tying each claim to the page it came from, and a follow-up like "something hands-on" means the agent writes the whole comparison again.

So people either accept an unsourced summary or open every page themselves to check it. Neither survives a second device or a reload, because nothing records which options were offered or which one was chosen.

The approach

The agent gets one new tool, present_choices. It prepares either clarification buttons or up to three sourced comparison cards, and a click comes back as a structured choice the agent continues from. The agent owns research and candidate copy. TypeSafe Jev makes two narrow decisions: whether to show the agent's prepared cards or answer in prose, and how well each candidate fits the user's latest message. Both go in one batched call. Code keeps control: Jev can only pick from the agent's control or prose, it overrides the agent only when its confidence is at least 0.5, and options are ranked by whole rubric level with ties kept in the agent's order, because fractional scores between levels are not calibrated finely enough to order.

Evidence before comparison. In live mode a comparison is refused unless every source URL was read successfully in the same turn, and each label, source title and detail appears in that page's text, ignoring case and spacing. A mail-based choice requires the thread to have been read. This costs the agent a browse per source, which is the point: a card cannot cite a page nobody opened.

The server owns panel state. Panels and selections are versioned records under jev_threads, updated by compare-and-swap, so stale, cross-thread and conflicting clicks are rejected and a failed selection can be retried. A new user turn retires the current panel before the agent runs. That matches the chat transcript, where any later user message makes earlier choices stale, even if the turn then fails or is cancelled. The retiring turn may still refine the panel, which is how "something hands-on" re-ranks the stored candidates without new browsing. A run that resumes after a tool result is not a new turn and keeps the panel.

Jev sees the person, not the agent's paraphrase. State carries the user's own latest message next to the agent's summary, and questions refer to options by position so page-derived labels stay out of the instructions.

Bounded provider cost. One TypeSafeClient is shared per server, with a 5-second timeout and one retry: about 10 seconds worst case instead of the SDK's default of 30 or more. Failures log only HTTP status and request ID. A 401 or other 4xx that cannot succeed tells the agent not to retry and to answer in prose.

JEV_MODE defaults to off. sample uses a scripted scorer and needs no key. live requires a server-only TYPESAFE_API_KEY and refuses to start without one. The same card component renders in web and native chat.

What is not covered

  • The live recording above predates the review fixes: the confidence gate, rubric-level ranking, user-message state and prompt wording. Rocky Shore is expected to stay first, but that has not been re-run with a live key.
  • Live mode sends the user's latest message, agent-written context (which can quote private mail) and every candidate's text and URLs to TypeSafe. The live-mode disclosure lists each field. Jev cannot approve or execute anything.
  • The sample scorer never picks prose, so the demo does not show that path. Service-level tests cover it.
  • Native bundles were exported, but no device or simulator was exercised.
  • The 0.5 confidence floor is a starting value, not tuned against real traffic.

Verification

Locally on the merged head: root, mobile and server TypeScript checks, Biome, the server build, and 275 of 275 tests. CI passes all seven jobs: quality, the web, iOS and Android builds, browser, browser-container and computer-container.

  • tests/jev.test.ts: the adapter against SDK-shaped answers, including the confidence gate, rubric-level ranking, positional references and failure logging. Contract tests drive the real TypeSafeClient through a fake fetch: request shape, one retry on 503, no retry on 401, abort.
  • tests/jev-persistence.test.ts: versioned panels, compare-and-swap selection, stale and cross-thread rejection, refinement and retry.
  • tests/conversation-jev.test.ts: evidence checks for mail, redirects, invented labels, titles and details. Also a panel followed by a failed or cancelled turn rejecting the old selection, a resumed run keeping it, and only the retiring turn refining it.
  • apps/mobile/test/jev-actions.test.ts: the transcript projection that decides which card accepts input.
  • tests/config.test.ts, tests/demo-model.test.ts: mode validation and the scripted demo.

@jerelvelarde
jerelvelarde marked this pull request as ready for review September 23, 2026 17:54

kvnloo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

I like that the server has a real durable head (jev_threads) and treats the panel/tool result as a projection of that state.

I think the mobile side may currently have a second source of truth, though.

latestJevPanelId() invalidates the visible panel as soon as it sees any later user message, while the durable server head is only expired by expireAfterOrdinaryTurn() when the run reaches RUN_FINISHED.

So this sequence looks possible:

  1. panel P is current;
  2. user sends an ordinary follow-up;
  3. the transcript immediately makes P stale in the UI;
  4. that run errors/stops before RUN_FINISHED;
  5. server jev_threads.currentPanelId is still P.

At that point the UI says “Earlier choices” and disables P, while JevService.select() still considers P current.

Could we add a regression for panel → ordinary user turn → failed run and make one layer authoritative?

Either the durable head should be expired when that new turn is admitted, or the UI should project panel availability from durable Jev state rather than independently reconstructing it from transcript order.

That seems especially important for Rich Threads / cross-device replay: ephemeral controls should be a projection of the durable interaction state, not a competing state machine.

AI-use note: I used an AI assistant to trace the Jev head lifecycle and mobile transcript projection and draft this review; I verified the described paths against the current PR head before posting.

kvnloo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

One integration-boundary question for live mode, especially now that this head includes a real live-Jev recording: could we document exactly what leaves OpenMuse when JEV_MODE=live?

LiveJevAdapter sends TypeSafe/Jev:

  • the user's message;
  • the model-prepared context;
  • the current selectedId;
  • the full candidate options, including labels, details, and source URLs.

That context can be derived from mail or other private workspace data even though Jev itself only ranks/selects and never gets authority to act.

The docs are now clear about which recordings use real Jev and that TYPESAFE_API_KEY stays server-side, but I couldn't find the corresponding operator disclosure for the data sent to the live service.

Could we add a short live-mode privacy/config note? Something along the lines of: enabling live Jev sends the request/context and prepared candidate data to the configured Jev provider for control selection/scoring; source data remains evidence and Jev cannot approve or execute actions.

I think that would make the otherwise nice authority boundary here explicit to operators.

AI-use note: I used an AI assistant to inspect the live adapter payload and current documentation and draft this review; I verified the listed fields against the current PR head before posting.

@jerelvelarde

Copy link
Copy Markdown
Collaborator Author

Review: Jev usage

Checked against the @typesafe-ai/sdk@0.6.0 type declarations and TypeSafe's docs (Models, Confidence, the re-ranking cookbook, and the jev-1.13 known-issues page).

Verdict: the SDK calls are correct and the integration keeps code in control. The gaps are in how the answers are used.

What's right

  • One batched systemOne call for the control choice and every fit score. TypeSafe recommends this batching.
  • Jev can only pick from [agent's control, "agent"], and every answer is re-validated.
  • The abort signal is passed through RequestOptions. The key stays on the server, and JEV_MODE=live won't start without one.
  • The model is pinned to jev-1.13.0, as TypeSafe recommends. The SDK default is the moving jev-latest alias.

Should fix

  1. confidence is never read. Low-confidence answers are acted on the same as certain ones. TypeSafe's docs say not to act on low confidence. Gate the control choice and fall back to prose when it's uncertain.
  2. Ranking sorts on fractional score values. The jev-1.13 known-issues page says values between levels are poorly calibrated, so 2.01 vs 1.99 currently decides the top-3 cut. Treat near-equal scores as ties and keep the agent's order.
  3. Jev scores the agent's paraphrase (message) instead of the user's latest message. The server already has the real message. Pass it in so Jev judges independently of the model it's checking.
  4. Failure causes are dropped. A 401 or 429 is never logged, and the agent is told to "retry" an auth failure.

Worth tightening

  1. A new TypeSafeClient is built on every run. Build one and reuse it.
  2. SDK defaults (10 s × 3 attempts, no total budget) can block present_choices for 30 s or more. Set timeout and retry explicitly.
  3. Web-derived labels are placed inside question instructions. Refer to options by position in state (options[0]) instead.
  4. The model default is written in two places.
  5. Test mocks (as never) leave out confidence, probabilities, model and usage. Add a contract test with a real client and a fake fetch.

Also: the description says "Jev chooses the prepared control", but the agent chooses the control and Jev decides whether to show it. Suggested wording: "Jev decides whether to show the agent's prepared choices and ranks them."

I'll push fixes for 1–9 to this branch.

- Only let Jev override the agent's prepared control when its choice
  confidence is at least 0.5; an uncertain answer keeps the agent's control.
- Rank options by rounded rubric level with stable ties, since fractional
  scores between levels are not calibrated finely enough to order.
- Send the person's latest message as `userMessage` in state, alongside the
  agent's summary, so Jev judges independently of the agent's framing.
- Refer to options by position (`options[i]`) in question text and keep
  page-derived labels in state only.
- Log provider failures with status and request ID, and tell the agent not
  to retry non-retryable 4xx responses such as 401.
- Build the adapter once per server so live mode shares one TypeSafe client,
  with a 5s timeout and one retry instead of the SDK's ~30s worst case.
- Keep the default Jev model in one place (config.ts).
- Test against SDK-shaped answers and add contract tests that drive the
  real TypeSafeClient through a fake fetch (request shape, retry, 401, abort).
@jerelvelarde

Copy link
Copy Markdown
Collaborator Author

Pushed cec18d7, which addresses 1–9 from the review above.

  1. Confidence. Jev now overrides the agent's prepared control only when its choice confidence is at least 0.5. One change from what I wrote above: an uncertain answer keeps the agent's control instead of falling back to prose. The agent already chose to present choices, so a low-confidence veto shouldn't discard that work.
  2. Ranking compares scores rounded to the nearest rubric level, with ties kept in the agent's order.
  3. User message. State now carries userMessage (the person's latest message, taken from the run) alongside agentSummary, and both questions refer to userMessage.
  4. Failures are logged with status and requestId, without bodies or keys. Non-retryable 4xx responses (401, 403, 400, 422) tell the agent not to retry and to answer in prose.
  5. The adapter is built once per server on first use, so live mode shares one TypeSafeClient.
  6. timeout: 5000 and maxRetries: 1, about 10 s worst case.
  7. Fit questions refer to `options[i]`. Labels stay in state, and a test checks that injected label text doesn't reach the instructions.
  8. The default model is defaultJevModel in config.ts. LiveJevAdapter now requires a model.
  9. Mocks now use the full SDK response shape. New contract tests drive the real TypeSafeClient through a fake fetch: request path, auth header, model, one retry on 503, no retry on 401, and abort. A conversation test checks that the user's typed text, not the agent's summary, reaches Jev.

Checks: pnpm typecheck, pnpm lint, pnpm build:server, and pnpm test (222/222) pass locally.

Not re-run: the live TypeSafe recording. Ranking now rounds to rubric levels and the prompt wording changed, so the Rocky Shore result in the recording should be re-checked with a real key. The sample adapter still never returns "agent"; that path is covered by the service-level tests in jev-persistence.test.ts.

The mobile transcript marks a panel stale as soon as a later user message
appears, but the server only expired the head on RUN_FINISHED. A follow-up
that failed or was cancelled left the UI showing "Earlier choices" while
select() still accepted the panel.

- Expire the current panel when a new user turn is admitted, before the
  agent runs, so the durable head matches the transcript on every device.
  Runs resuming after a tool result are not new turns and keep the panel.
- Record which turn retired the panel so that turn can still refine it
  ("Something hands-on"); a later turn cannot, and selection stays rejected.
- Regression tests: panel -> ordinary turn -> failed run, and -> cancelled
  run, both reject the old selection; resumed runs keep the panel; only
  the retiring turn may refine.
- Document exactly what JEV_MODE=live sends to TypeSafe and what Jev can
  and cannot do, with pointers from SECURITY.md and .env.example.
@jerelvelarde

Copy link
Copy Markdown
Collaborator Author

@kvnloo thanks, both addressed in 97dfa78.

Two sources of truth. The server is now authoritative, and it agrees with the transcript. The current panel is retired when a new user turn is admitted, before the agent runs, instead of on RUN_FINISHED. So a follow-up that fails or is cancelled leaves the server head at null, matching what latestJevPanelId() already shows, and select() rejects the old panel. Two details:

  • A run that resumes after a tool result (last message isn't a user message) isn't a new turn, so it keeps the panel.
  • Refinement is itself a new user turn ("Something hands-on"). The head therefore records which turn retired the panel, and only that turn may refine it. A later turn can't, and selecting it stays rejected.

Regressions in tests/conversation-jev.test.ts: panel → ordinary turn → failed run, and → cancelled run (both check currentPanelId === null, that replaying the selection gives RUN_ERROR, and that latestJevPanelId agrees), plus a resumed run keeping the panel and a later turn failing to refine. The old "cancelling keeps the choice" test encoded the divergence, so I replaced it.

I kept the transcript projection on mobile rather than adding a durable-state endpoint. With the server matching it, both stay consistent across devices and replay.

Live-mode disclosure. New section What live mode sends to TypeSafe. It lists each field (userMessage, agentSummary, context which may quote private mail, selectedId, and full candidate labels, details and source URLs), what isn't sent, and that Jev only picks cards-or-prose and scores options: it can't approve or execute anything. Failures log only status and request ID. SECURITY.md and .env.example point to it.

Locally: typecheck, lint, server build, and 225/225 tests pass.

Resolve package.json and pnpm-lock.yaml: keep main's TanStack AI
dependencies (CopilotKit#46) and add @typesafe-ai/sdk.
Resolve tests/config.test.ts: keep both the Jev mode test and main's
shadowed-.env test (CopilotKit#57).
@jerelvelarde jerelvelarde changed the title Add Jev-guided choices and sourced comparison cards to OpenMuse Answer ambiguous requests with choices and sourced comparison cards Sep 29, 2026
@jerelvelarde
jerelvelarde merged commit 4f3f7c7 into CopilotKit:main Sep 29, 2026
7 checks passed
sunshaoan0808 pushed a commit to sunshaoan0808/openmuse that referenced this pull request Sep 30, 2026
上游 5 个提交:JEV(CopilotKit#42)、无 scheme worker 地址(CopilotKit#92)、web 输入框焦点(CopilotKit#89)、
worker 测试健壮性(CopilotKit#90)、render 蓝图(CopilotKit#86)。

冲突解决(4 文件 9 块)与本地适配:
- conversation.ts:以上游 JEV 结构(runInternal/choices 分支/noteEvidence)为基座,
  换回我们的中性工具层(forAgUi + chatTools),并把 noteEvidence 改成经 ChatToolContext 回调;
  prompt 用我们的 chatInstructions() + JEV 段 + computerInstructions;恢复 maxSteps=10 与
  sample 分支的 threadId 注入
- agent.ts:保留我们的 mastra 引擎分支,两处 ConversationAgent 注入 sharedJevAdapter()
- chat.tsx:我们的卡片(搜索/浏览器操作/审批)与 JEV 卡片并存,去重 delegate_task 渲染器
- config.ts:去掉 cherry-pick 造成的 browserWorkerUrl 重复定义
- mastra-engine.ts:JEV 工具与指令补到 mastra 路径(上游只挂自带引擎,我们线上跑 mastra)
- jev/tools.ts:抽出 presentChoicesSpec 中性规格,两条引擎共用同一份校验逻辑

验证:pnpm test 321/321 通过(含上游 JEV 全部测试)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants