Skip to content

Windows: roam, multi-model chat, island questions, hats and lighter memory - #109

Open
arnavaggarwal-dev wants to merge 14 commits into
Louis-CFM:mainfrom
arnavaggarwal-dev:windows-roam-chat-hats
Open

arnavaggarwal-dev wants to merge 14 commits into
Louis-CFM:mainfrom
arnavaggarwal-dev:windows-roam-chat-hats

Conversation

@arnavaggarwal-dev

@arnavaggarwal-dev arnavaggarwal-dev commented Oct 2, 2026 •

Copy link
Copy Markdown

Windows-side features I've been running daily. Everything is under windows/; macOS is untouched.

What's in it

  • Roam: drag Mochi out of the island to look at the screen. It scans, screenshots, and opens the chat with the image attached. Cancel with Esc, a right click, or by dropping it back on the island. A "quick screenshot" setting skips the scan animation. On Linux the overlay is a transparent layer-shell surface that takes the keyboard (Esc) but no clicks; the island forwards the held pointer, and the screenshot uses the desktop's tool (spectacle, gnome-screenshot, grim, or import/scrot on X11).
  • Area scan: hold Shift while carrying Mochi to pick an area instead of the whole screen. The rest of the screen dims, the area shows its size, and on drop Mochi hops onto its edge, scans only inside it, and only that area is captured. Windows captures the area directly; on Linux it is cropped out of the full screenshot (shot.rs, with a unit test).
  • Chat providers: any OpenAI-compatible endpoint (OpenRouter, NVIDIA NIM, Groq...), a saved model list, and a picker next to Send. Keys live in Credential Manager, one per host. Replies render GitHub-style Markdown (below). Sending an image to a model that can't take images asks first.
  • Questions on the island: AskUserQuestion multiple choice is answerable from the card, labelled yes/no, pick one or pick any. The card closes live when answered in the terminal. Claude Code's questions open even when another Mochi is focused.
  • Mochi names, shapes and hats: each Mochi can be renamed and given a body shape and a hat. The vector art is generated by windows/scripts/wardrobe. Hats ride the head on a fixed seat (no drifting); only tassels swing.
  • GitHub card: latest Actions runs with a bar of finished steps while one runs, latest pushes, stars and repos in the header; refreshes every 20 s during a run. A git push that Claude Code runs pops the card open with the push in progress (from the hooks, since GitHub can't see a push in flight), keeps it open until the push ends, then shows Pushed or Push failed.
  • Chat extras: replies end with a hidden {"mood": ...} that Mochi acts out (new sad, shy and scared emotes); code blocks have a copy button; a dropped code file is sent as a fenced block and shown as a code card; "Clear" under Mochi, which also drops the attached file so it isn't sent again.
  • Upload: the bar never moved (nothing updated State.uploadProgress); it now follows Mochi's progress curve and says the file is read locally, not uploaded. The Mochi that eats the file is the focused one, in its colour, shape and hat, and the bar and glow take its colour.
  • Closed island: the middle shows the focused Mochi's current step, or the time when nothing is going on.
  • Image, video, speech and 3D models: a saved model says what it makes (guessed from its id: flux → images, sora → video, orpheus → speech, trellis → 3D, whisper → transcribes speech).
    • Images use /images/generations, or the chat route with modalities for OpenRouter. While one is made, the chat shows a glowing tile with a sparkle and the seconds.
    • Video uses OpenAI's Videos API (create, poll with progress, download), with a progress bar; it gives up after 20 minutes.
    • Speech uses /audio/speech; Groq's Orpheus gets its six voices, and its 200-character limit is checked before sending.
    • 3D uses Pollinations' /3d route (TRELLIS 2, low/medium/high detail). The dropped picture is uploaded to Pollinations' media store first, and the chat says so. The reply shows a rendered still; "View in 3D" opens a turntable. three.js is loaded only when a model is shown.
    • Results show inline with Download (to Downloads, never overwriting) and full screen on the roam overlay (Esc, click outside or close). The webview can only read a generated file by its bare name in the media folder, and the API key only goes to the provider's own host.
  • Mic: mic and Send are one button, a mic while the field is empty and Send once it holds text.
    • Words appear while you talk. Speech is split into phrases at short pauses and sent to your saved speech-to-text model (e.g. Groq Whisper), with a long phrase re-read about every 1.5 s, under a budget of about 18 requests a minute.
    • It stops and sends after 1.6 s of quiet. Silence is never uploaded, and Whisper's stock inventions ("Thanks for watching", a lone "Thank you." from noise) are dropped. Esc cancels; leaving the chat or hiding the island releases the mic.
    • The mic turns off after each prompt; Settings → General → "Keep the mic on" listens again after the reply to a spoken prompt.
    • On Windows the island grants its own page the microphone. Linux isn't wired yet (WebKitGTK needs a permission handler).
  • Markdown like GitHub: replies go through marked + marked-footnote.
    • Tables with alignment wrap to fit the chat instead of scrolling sideways. Also task lists, strikethrough, autolinks, nested lists, footnotes, details/summary, sub/sup/kbd, and GitHub's alerts (> [!NOTE] and friends).
    • Math ($…$, $$…$$, ```math, \(…\); "$5 and $10" stays text) is typeset by KaTeX, loaded the first time a reply has math. Its fonts are emitted as files, since the CSP has no font-src data:.
    • Model output is untrusted and the page holds Tauri IPC. So the HTML goes through DOMPurify with an allowlist: no script, on*, style or class, and ids are prefixed user-content-. Images never get a live src, and no link keeps an href: web links open in the browser through Rust.
  • Fixes:
    • File drops reach the island at last. WebView2's own child window (owned by its browser process) is what Windows asks during a drag, so Tauri's native drop events never fired, and no app can revoke another process's drop target. The island now takes a plain HTML5 drop and sends the bytes to a new ingest_bytes command; the same code serves Linux.
    • An integration's key button no longer saves an empty key over a stored one: it becomes "Clear key" and asks first.
    • The island is a tool window, so taskbar replacements don't list it as an app.
    • Tool step labels are in English (they were the macOS French, e.g. "Exécute").
    • Sweat drops run down instead of floating up like the other particles.
    • Settings no longer showed a newly added model twice (two overlapping redraws each added their rows).
    • Linux build: roam_linux.rs imports tauri::Manager (thanks @arreina).
  • Music (Windows): a Music pill (Melody) that reads the Windows media controls, the same feed as the volume flyout's player, so Spotify, Apple Music, YouTube Music and any player reporting to it show the track with play/pause and skip. Event-driven (GlobalSystemMediaTransportControlsSessionManager events), so nothing polls. Mochi dances only to music: browser sessions (YouTube videos) don't count, installed web apps like YouTube Music do. While dancing, the eyes groove to the beat instead of following the cursor, music notes fly off, and the head opens like a real AirPods case (lid hinged at the back, white inside, two wells, front light): the buds rise out and shoot off to the left and right edges of the monitor on the screen overlay (borrowed like the media preview, click-through, skipped when busy), then the lid snaps shut. 15 s after the music stops the lid opens, the buds swoop back in and drop into their wells. The buds are drawn in code (airpod.ts). Linux reads nothing yet (MPRIS would be the route).
  • Branch updated with main (0.1.3). One merge commit, no conflicts.
  • Music, steadier: skipping no longer flashes "Nothing playing" (the last track holds for 3 s while the player reports nothing between songs), and Melody takes the colour of the song's cover (pill, card and its Mochi; red without art).
  • Ticker fix: Claude's tool activity froze after a session's 20th step, because the capped step list stopped its index moving. It follows a running step count now. With several sessions, one session's Stop no longer idles Mochi while another works. (The macOS ticker has the same cap issue; untouched.)
  • Loaders: speech plays a music-note Lottie recoloured with a turning rainbow (lottie-web light canvas build, loaded only while speech generates); image and video show an "aurora" placeholder instead of the sparkle tile.
  • Text: no em dashes in on-screen strings.
  • Smaller: "keep island visible", auto-close down to 3 s, single-line ticker.
  • Memory: the three webviews share one renderer process, and hidden ones ask WebView2 to trim. Working set dropped from about 790 MB to 650 MB on my machine.

Heads-up

  • Some of these change existing views (ticker, question card height, settings sections, the closed island), which CLAUDE.md asks to avoid unless requested. Happy to split those out or drop them.
  • The first commit is large. I can break it into smaller PRs per feature if that's easier to review.
  • New dependencies, against CLAUDE.md's "no third-party dependencies unless truly unavoidable": three.js (showing 3D models), marked + marked-footnote (GitHub-flavoured Markdown), DOMPurify (sanitizing it) and KaTeX (math), and lottie-web for the speech loader. The heavy ones (three.js, KaTeX) load only when a reply needs them. Writing a safe GitHub-compatible parser, sanitizer or 3D renderer by hand didn't seem wise; happy to discuss.

Tested

  • Windows 11: cargo test (29 pass), tsc --noEmit, release build, daily use. The area scan's animation and capture call were also checked in a headless harness.
  • Markdown: src/dev/markdown-check.ts (33 checks in headless Edge, including hostile HTML: <script>, <img onerror>, javascript: links, style, id clobbering, KaTeX \href). Dictation: src/dev/dictation-check.ts with a fake microphone and stubbed transcription. The 3D viewer was checked by rendering Khronos's sample Duck.glb.
  • Music: checked live with YouTube Music (app) and Chrome tabs; music.rs unit tests cover browser vs. player detection and names. The lid and buds were checked rendered in headless Edge.
  • Not yet tried against the real providers (image, video, speech, 3D and transcription calls), or with a real microphone.
  • Linux: not built locally. The roam overlay, area crop, screenshot tools and HTML5 drop are written for it but untested; the Linux CI run on this PR should cover the build.

…emory

Drag Mochi out of the island to look at the screen: it scans, takes a
screenshot and opens the chat with it attached. Cancel with Esc, a right
click or by dropping it back on the island; a "quick screenshot" setting
skips the scan. Holding it too long makes it dizzy and it flies home.

Chat: custom OpenAI-compatible providers (OpenRouter, NVIDIA NIM...) with a
saved model list and a picker next to Send. Keys stay in Credential Manager,
one per host. Replies render Markdown. Sending an image to a model that
can't read images asks first.

Island: AskUserQuestion multiple choice can be answered from the card, with
a label for yes/no, pick one or pick any. The card closes live when the
question is answered in the terminal. Claude Code's questions always open
even when another Mochi is focused.

Mochis can be renamed, and each can wear a body shape and a hat (vector
art generated by scripts/wardrobe). Hats ride the head on a fixed seat, so
they never drift; only tassels swing.

Other: dropped files reach the island again (WebView2's own drop target is
revoked), "keep island visible", auto-close down to 3 s, a single-line
ticker. The island is a tool window, so taskbar replacements don't list it.

Memory: the island, settings and roam windows share one renderer process,
and hidden windows ask WebView2 to trim (about 790 MB down to 650 MB of
working set here).

Roam needs Win32 capture, so other platforms get stub commands that hand
Mochi straight back.
File drops never reached the island on Windows. During a drag, Windows asks
the window under the cursor, which is WebView2's own child window, owned by
its browser process, and with dragDropEnabled it refuses everything. Tauri's
handler sits on our parent window and is never consulted, and the old guard
couldn't revoke a drop target that belongs to another process. The island
now takes a plain HTML5 drop (dragDropEnabled off) and sends the file's
bytes to a new ingest_bytes command, which saves into the same inbox. The
same code serves WebKitGTK on Linux. The guard and unblock_webview_drops go.

Roam on Linux: the overlay is a transparent layer-shell surface over the
whole monitor that takes the keyboard (Esc) but no clicks. While the button
is held the compositor keeps sending the pointer to the island, so the island
forwards it until the release drops Mochi. The screenshot comes from the
desktop's own tool: spectacle, gnome-screenshot, grim, or import/scrot on X11.

Settings: with a key stored and the field empty, the integration key button
used to save an empty key over it. It now reads "Clear key" and asks first;
removing the Claude key asks too.
GitHub card: your latest Actions runs, with a bar of finished steps while one
is running, and your latest pushes; stars and repos move into the header.
It refreshes every 20 s while a run is going and every 5 minutes otherwise,
and a finished run is a pill event.

Chat: the model ends each reply with {"mood": ...}, which is taken off before
display and acted out by Mochi (new sad, shy and scared emotes; celebration
is the hop). Code blocks get a language label and a copy button. A dropped
text or code file goes to the model as a fenced Markdown block and shows in
the chat as a code card. "Clear" under Mochi starts the conversation over.

Upload: the bar and percentage never moved because nothing updated
State.uploadProgress after resetting it; they now follow Mochi's eased
progress curve. The label says what happens: the file is read into the
local inbox, nothing is uploaded.
@arreina

arreina commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Heads-up from a Linux build check (cargo check --workspace --all-targets on Linux Mint 22.3, with this PR merged onto current main, head 75db4c7): it doesn't compile on Linux.

roam_linux.rs:81 and :131 call app.get_webview_window(…), which needs the Manager trait in scope (E0599). Adding use tauri::Manager; at the top of roam_linux.rs fixes it: I checked, and with that one line the whole workspace builds on Linux with no errors.

(Linux CI now runs on pull requests that touch windows/, so this will show up there too once it runs.)

GitHub: a `git push` Claude Code runs pops the GitHub card open with the push
in progress, held open until it ends (the island's auto-close would fold it
away mid-push), then shows Pushed or Push failed and refreshes from GitHub.
Pushes come from each repo's pushed_at and the commits API, since the
activity feed runs minutes behind and no longer carries commits.

Upload: the eating Mochi is the focused one, in its colour, shape and hat,
and the bar, glow, border and crumbs take its colour; the label names it.

Closed island: the middle shows the focused Mochi's current step, or the
time when nothing is going on.

Also: tool step labels in English (they were the macOS French), sweat beads
run down instead of floating up, and Clear chat also drops the attached file.
Holding Shift while carrying Mochi pins where it is and stretches an area
from there to the cursor: the rest of the screen dims, the area gets
marching dashes, corner brackets and its size. Dropped, Mochi hops onto the
area's edge, a grid and one sweep run inside it only, a pulse locks it in,
and only that area is captured and flashes. Letting go of Shift before the
drop goes back to the whole-screen scan.

Windows captures the area directly. Linux screenshot tools only take the
whole screen, so the area is cropped out of it (png crate, already in the
lockfile through Tauri), with a unit test.

The README's controls table now lists dragging Mochi out, and Shift.
get_webview_window comes from the Manager trait; without it the Linux build
fails with E0599. Windows never compiles this file, so it went unnoticed.
@arnavaggarwal-dev

Copy link
Copy Markdown
Author

Thanks for checking on Linux! Fixed in ba196cf: roam_linux.rs now imports tauri::Manager. The PR has also moved on since 75db4c7 (Shift-drag to scan an area, which crops the screenshot on Linux), so the Linux CI run on this push covers that too.

Each saved model showed its name in the name box and again in the line
under it, since the name defaults to the model id. The line now names the
provider, plus the model id only once the model has a name of its own.
A saved model now says what it makes: text (the chat, unchanged), images,
video, speech, 3D, or transcribed speech for the mic. It's guessed from the
model id when one is added (flux -> images, sora -> video, orpheus -> speech,
trellis -> 3D, whisper -> transcribes) and can be changed on each row.

- Images: /images/generations (b64_json or url), or the chat route with
  modalities for OpenRouter. While one is made the chat shows a glowing
  rainbow-rimmed tile with a sparkle and the seconds.
- Video: OpenAI's Videos API (multipart create, poll every 4 s with
  progress, then /content), with a bar and percentage; gives up at 20 min.
- Speech: /audio/speech. Groq's Orpheus gets its six voices and its
  200-character limit is checked before sending.
- 3D: Pollinations' /3d route (TRELLIS 2 by default, low/medium/high
  detail). The dropped picture, or the last image made, is uploaded to
  Pollinations' media store first and the chat says so. The chat shows a
  rendered still; "View in 3D" opens a turntable. three.js is loaded only
  when a model is shown, and WebGL contexts are freed right after.
- Mic: right of Send when a speech-to-text model is saved. It records,
  shows the words as they're heard (every 4 s, within free-tier limits) and
  sends on a second click or Enter; Esc cancels. The island grants its own
  page the microphone (WebView2 permission request).
- Results: inline image, video or audio with Download (to Downloads, never
  overwriting) and full screen on the roam overlay, which the overlay
  refuses while Mochi is roaming. Esc, a click outside or the close button
  end it; on Linux the overlay keeps its input region across remaps.

Generated files live in local_dir()/media; the webview can only read, copy
or preview a file there by its bare name, and the API key only goes to the
provider's own host.

Also: Settings no longer shows a newly added model twice (two overlapping
redraws each appended their rows).
Mic and Send are one button: a mic while the field is empty (and a
speech-to-text model is saved), Send as soon as it holds text.
- Dictation listens in the page and splits speech into phrases at short
  pauses (550 ms). Each phrase goes to the saved speech-to-text model as a
  16 kHz WAV; a long phrase is re-read about every 1.5 s so words show up
  while you're still talking. A request budget keeps it under ~18 a minute.
- Speech is louder than a fixed floor and 3x the learned room noise; under
  250 ms of voice is never uploaded, so silence never leaves the PC.
  Whisper's stock inventions ("Thanks for watching", a lone "Thank you."
  from noise) are dropped.
- It stops and sends after 1.6 s of quiet, gives up after 8 s of no speech,
  and caps at 2 minutes. Esc cancels (even while transcribing); click or
  Enter sends now. Leaving the chat, collapsing the island or hiding the
  page releases the mic.
- The mic turns off after each prompt. Settings > General > "Keep the mic
  on" (off by default) listens again after the reply to a spoken prompt,
  and stops after 30 s of silence.

Replies render GitHub-flavoured Markdown through marked + marked-footnote:
tables (aligned, wrapping to fit the chat instead of scrolling), task lists,
strikethrough, autolinks, nested lists, footnotes, details/summary,
sub/sup/kbd, and GitHub's alerts (> [!NOTE] and friends). Math ($...$,
$$...$$, ```math, \(...\), \[...\]; "$5 and $10" stays text) is typeset by
KaTeX, loaded only the first time a reply has math; its fonts are emitted
as files because the CSP has no font-src data:.

Model output is untrusted and the page holds Tauri IPC, so marked's HTML
goes through DOMPurify with an allowlist (no script, on*, style or class;
ids prefixed user-content- against clobbering), images never get a live
src (they become links), and no link keeps an href: http(s) ones open in
the browser through Rust. KaTeX runs with trust off and a macro cap.

src/dev/markdown-check.ts (33 checks incl. hostile HTML) and
src/dev/dictation-check.ts (fake mic, stubbed transcription) are dev-only
self-checks.
The Music pill reads the Windows media controls (the volume flyout's
player feed), so Spotify, Apple Music, YouTube Music and any player that
reports to it show up with play/pause and skip. Event-driven: nothing
runs until a session changes.

Only music makes Mochi dance: browsers' sessions (YouTube videos) don't
count, installed web apps such as YouTube Music do. While dancing, the
eyes groove to the beat instead of following the cursor, music notes
fly off, and the head opens like an AirPods case: two buds fly out
beside the island and bob to the beat, then fly home and the lid shuts
15 seconds after the music stops. The music card now refreshes when the
track changes.
The head opens like a real AirPods case: the lid tips up and back on a
hinge at the back of the seam, seen a little from above, showing its
white underside, the two wells and the buds standing in them, with the
front light glowing. The buds rise out, shoot off to the left and right
edges of the monitor on the screen overlay (spinning, growing, with a
motion trail and sparkles), and the lid snaps shut. Fifteen seconds
after the music stops the lid opens, the buds swoop back in from the
edges, drop into their wells and the lid shuts on them.

The buds are drawn in code (airpod.ts), shared by the case and the
overlay. The flight borrows the roam overlay like the media preview,
click-through, and is skipped when the overlay is busy. The 15 s wait
runs on a timer, so nothing keeps the frame loop going meanwhile.
The buds grow to 190 px on their way across the screen (was 76).
… loaders

- Music: skipping a track no longer flashes "Nothing playing". When the
  player briefly reports nothing between tracks, the last track stays on
  show for 3 s (a one-off timer, no polling), the controls keep working
  and Mochi keeps dancing. The "Changing" status counts as playing.
- Melody takes the colour of the song's cover: Windows' image decoder
  scales the art to 16x16 and the busiest vivid hue wins, lifted until
  it reads on the black island. It colours the pill, the card and the
  integration Mochi, and falls back to the red without art.
- Ticker: Claude's tool activity froze after a session's 20th step (the
  step list is capped at 20, so its index stopped moving and the ticker
  never scrolled again). It now follows a running step count. With
  several sessions, one session's Stop no longer idles Mochi while
  another is working.
- Loaders: speech plays a music-note Lottie recoloured with a turning
  rainbow (lottie-web's light canvas build, loaded only then); image and
  video show an "aurora", the coming picture's shape with colour
  drifting under a band of pixels developing.
- No em dashes in on-screen text.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants