Skip to content

Chat with local models (Ollama, LM Studio, OpenAI-compatible) and Markdown answers (Tauri build) - #173

Open
AinzDerErste wants to merge 14 commits into
Louis-CFM:mainfrom
AinzDerErste:local-models
Open

AinzDerErste wants to merge 14 commits into
Louis-CFM:mainfrom
AinzDerErste:local-models

Conversation

@AinzDerErste

@AinzDerErste AinzDerErste commented Oct 3, 2026 •

Copy link
Copy Markdown

Part 3 of 3. Draft until #158 is merged: it sits on top of #158 and of the two PRs before it (#171, #172), so until they are merged the diff also shows their commits. Only the last commit, "Chat with local models, and Markdown in the answers", is new here.

#156 for the Tauri build, plus any OpenAI-compatible server:

  • Settings → Local models: Ollama and LM Studio (empty address = the usual one on this machine), and "OpenAI-compatible" for any other server that speaks the OpenAI API (I use an Unsloth server) with an optional key kept in the keychain, never in the settings file. Connect checks that the server answers and lists its chat models (embedding and rerank models left out); "Chat with" switches the chat to it; Disconnect forgets the address and the key.
  • The answer is streamed (POST /v1/chat/completions) and shown as it is written, at most 15 times a second; <think> blocks stay hidden while open and are dropped from the finished answer. A dropped text file goes along inline (24 000 characters at most), images and PDFs by name only. The history is kept apart from Claude's, and a local model is not told it has web search.
  • Markdown in the chat's answers for every provider, as on the Mac: headings, lists, bold, italic, inline code, links (only http/https), quotes, rules, and code blocks with a copy button. It is built as DOM nodes, never HTML. The system prompt now asks for light Markdown instead of forbidding it, as in Chat with local models through Ollama and LM Studio, and Markdown in answers #156.

A streamed answer

The picture comes from the slowed-down tests/fake_local_llm.py of this repo, so the answer is the fake one.

Checked: tsc, vite build, cargo test (URL cleaning, event stream parsing, think filtering, model list and connect against a one-shot local HTTP server, the key header, file inlining), no new warnings; a throwaway test of the Markdown parser and the link rule. The streaming against the fake server, and real chats with my own Unsloth server on KDE/Wayland, with a code block copied out of an answer.
Not checked: real Ollama and LM Studio servers, Windows, macOS, GNOME.

Not ported: the Gemini and OpenAI cloud providers of the Mac chat. The French docs still say local models are macOS only; I left them alone and will change them if you want.

Justin Minkmar added 14 commits October 3, 2026 13:26
With NVIDIA's proprietary driver, WebKitGTK's DMABUF renderer commits frames to a
Wayland surface after announcing explicit sync without an acquire point. The
compositor then closes the connection with a protocol error and the app dies on its
first frame.

Switch the DMABUF renderer off when the NVIDIA driver is loaded, unless the user set
WEBKIT_DISABLE_DMABUF_RENDERER themselves. Split prepare_environment into one
function per concern.
Wayland has no main display, and the layer surface was never assigned an output,
so the compositor picked one.

- Pin the layer surface to the chosen monitor (gtk_layer_set_monitor).
- Settings -> Island lives on lists every display by name and size.
- "Main display" now asks XWayland's RandR for the primary output, which KWin and
  Mutter set from the user's choice, and falls back to the first display.
- Report the compositor's pointer leave (GTK leave-notify) to the page. WebKitGTK
  sends no mouseleave when the pointer leaves the surface, so the island never
  learned the mouse was gone and never auto-closed.
- Restart the collapse timer when an alert reaches an island that is already open.
  forceHome cancelled it and only a state change started it again.
- Start the drop sequence after the island has opened. Waking a closed island passes
  through its default view, which stopped the sequence and left "uploading" stuck.
- Let clicks through #content and the hidden DOM view to the drop canvas' buttons;
  Cancel and Ask were covered by invisible elements.
- Give the ticker's dim text top:0. Without it WebKit places the absolute box one
  line below the text it should overlap.
- Let notes dismiss themselves after five seconds.
…there is no editor

- The relay passes Konsole's D-Bus names. The app switches to the tab over D-Bus
  and raises its window with a short KWin script, found by Konsole's process id and
  the tab's title.
- Look for code, code-insiders, code-oss, codium, vscodium and cursor in that order.
- When no editor is installed the island says so instead of silently opening the
  file manager.
- open_in_vscode returns an enum instead of a bool.
Settings -> Language (system, English, Deutsch, Français). An explicit choice tells
the chat to always answer in that language, whatever the question or the search
results are in. The step labels ("Runs", "Führt aus", "Exécute") follow it, where
they were fixed to French.
Without an API key the chat runs `claude -p` on the user's own login instead of
failing with "API key missing". Read, WebSearch and WebFetch are pre-approved, the
user's hooks are left out (--setting-sources project, run from a temp directory) so
the helper does not show up as a session, and follow-up questions resume the same
session.
Embed the handful of Font Awesome Free 7 (solid) paths in use, with the CC BY 4.0
notice, and drop the hand-drawn header icons. The tabs get a little more room.
Clicking the overview's left card opens an editor view while the session has
edited a file or run a command: Mochi and the Read / Edit / Bash / Done phases on
the left, on the right the changed lines in red and green with a few lines of
context, and under them the last command and what it printed. Nothing opens by
itself.

- snippet.rs reads the lines around an edit, only from a regular file inside the
  session's folder, up to 2 MB.
- The relay forwards the last three lines a Bash command printed, without colour
  codes; the rest of tool_response stays out.
CSS cannot transition between two gradients, so the halo's colour flipped in one
step. Ease colour and opacity per frame, like Mochi's position.
GTK's leave-notify reaches the page over IPC and can overtake the last mouse-move
WebKit was still handing it. That move put the mouse "inside" again after the leave,
and since nothing follows a pointer that is gone, the island never started its
auto-close timer; it stayed open until OK was pressed.

Ignore page mouse-moves for 400 ms after a leave. A real re-entry is picked up by the
next moves right after.
The view read its data from the one shared "Claude Code" pill. With several sessions
running, a new prompt in one of them wiped what the view showed and the project name
jumped between them.

Each session now has its own record, keyed by session_id, and the view shows the one
that was active last and has something to show. A session remembers the folder it
started in: its name, the paths in the diff and the folder the file is read from no
longer change when the shell cd's. A finished session is dropped after its view has
been shown; one that never says so is forgotten after 30 minutes.
As on the Mac (docs/INTEGRATIONS.md, "Jauge de forfait Claude"), for the Tauri build:

- A pill in the header of the overview ("Claude 73%", green below 50 %, orange up to
  80 %, red above, grey without data). Clicking it puts the plan card in place of the
  left card (5 hours and week, reset times, how old the numbers are), and Mochi wears
  the plan's colour while it is open. It closes when the view, the mode or the focus
  changes. Off by default.
- The numbers come from Claude Code's own status line (`rate_limits`), like on the Mac,
  and are checked the same way: 100-200 % is shown full, anything else is dropped, a
  reset more than 400 days away is milliseconds in disguise, and a window whose reset
  has passed counts as 0 %.
- The relay runs as the status line (`coucou-hook --statusline`), passes the limits on
  and runs the status line the user had before with the same input, printing its output
  untouched (10 s limit). That one is kept in statusline-previous.json beside the relay.
- Installing and removing the relay is its own step, apart from the hooks: Settings ->
  Plan usage shows the diff of the `statusLine` key only, backs settings.json up and
  writes after a click. Only `command` is swapped, so padding and the like stay;
  removing puts the previous status line back, or the key goes if there was none, and a
  status line the user changed since is left alone.
- Settings -> Plan usage -> Show in the notch turns it on and starts the install first
  when the relay is not in yet.

Calls using the older `coucou-hook StatusLine` form keep working.
Claude Code 2.1.85+ sends AskUserQuestion as a PreToolUse. The island now shows the
choices and hands the answer back, as the Mac app does (docs/INTEGRATIONS.md,
"Répondre aux questions").

- A dedicated hook entry (matcher `AskUserQuestion`, `coucou-hook --ask`, 130 s) is
  added by the installer next to the general PreToolUse one, which ignores that tool.
  Hooks installed before this show "Hooks outdated" in Settings with an update button;
  the usual diff, backup and confirmation apply.
- The relay waits up to 125 s and prints the documented PreToolUse output: allow, with
  `updatedInput` carrying the questions whole plus `answers`. Single-choice values are a
  string, multi-select an array. No answer, or an unusable one, prints nothing and
  Claude Code asks in the terminal. If Coucou is closed it exits at once, as before.
- One question at a time with a 1/N counter, the options as chips, "Other…" for free
  text, and "Reply in terminal". A single question with a single choice is answered by
  the click; several questions or a multi-select use Next / Send.
- The text field only takes the keyboard when it is clicked, so a question never steals
  what is being typed in the terminal.
- A question that can't be shown (not 1-4 questions, fewer than two options, a card
  already up, another agent, paused) is handed straight back to the terminal.
  PermissionRequest for this tool is declined the same way, for older Claude Code.
  Answering in the terminal takes the card away.
As the Mac app does since Louis-CFM#156 (docs/INTEGRATIONS.md, "Modèles locaux"), plus any
OpenAI-compatible server:

- Settings -> Local models connects Ollama or LM Studio (empty address = the usual one
  on this machine) or a server of the user's own that speaks the OpenAI API, with an
  optional key kept in the keychain. Connect checks the server answers and lists its
  chat models (embedding and rerank models are left out); "Chat with" picks who the
  chat talks to. Disconnect forgets the address, and the key.
- The answer is streamed (POST /v1/chat/completions) and shown as it is written, at
  most 15 times a second. <think> blocks of reasoning models stay hidden while open
  and are dropped from the finished answer. A dropped text file goes along inline,
  24 000 characters at most; images and PDFs by name only. The history is kept apart
  from Claude's, which carries tool blocks a local server would not understand.
- The chat asks for light Markdown now (the "no markdown" rule is gone from the system
  prompt, as on the Mac) and renders it for every provider: headings, lists, bold,
  italic, inline code, links, quotes, rules and code blocks with a copy button. It is
  built as DOM nodes, never HTML, and only http(s) links open.
- A local model is not told it has web search.

Nothing leaves the machine unless the address points elsewhere.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant