Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,24 @@ images to GHCR.
hosted-only release *and* proves the guard refuses the half-working one, and `k8s-smoke`
upgrades a live release into the mode and asserts the core answers. `core-app`
0.128.0→0.129.0 (MINOR), chart 0.1.2→0.2.0 (MINOR).
- **A deployment with no local AI says so once, instead of warning forever** (#962) — the web
shell had two inline strings for a dead local runtime ("The local runtime is unreachable — is
the ollama service up?", "local runtime unreachable") and rendered every other Ollama control —
the pull card, the catalog, the KV-cache card, the context-window card, the `Local (Ollama)`
optgroup — as though a runtime existed. On a deployment that deliberately runs **none** (hosted
chat, hosted embeddings, `OLLAMA_URL=""`) that reads as a fault the operator is supposed to fix,
when in fact nothing is wrong and there is nothing to fix. The shell now reads the runtime's
state from `GET /platform/v1/llm/local-runtime` — never from a failing model list, which since
this change answers `[]` with a 200 in *both* unhappy states and so carries no information at
all — and when that state is `absent` the local half of the Models page collapses into one
line: local AI is not configured on this deployment, hosted models are unaffected. The five
Ollama cards are removed rather than disabled, the embedding picker drops its local group so an
unrunnable model cannot be chosen, the chat picker drops its `Local` heading while keeping the
core-default row (that default may itself be hosted), the first-run welcome stops offering a
pull, and a module's model slot with nothing left to offer says why instead of looking broken.
`unreachable` keeps today's warning, because that state *is* an error and must still look like
one, and an older core with no such endpoint keeps today's behaviour exactly.
`web` 0.150.0→0.151.0 (MINOR).
- **The bound on a turn is the operator's; runaway is caught by behaviour** (#925) — the
**Agent cycles** setting stopped at 12, and the route enforced it *silently*: type 40 and 12 was
stored. A genuinely long task — search → read → read → summarize → write — ran out of rounds and
Expand Down
32 changes: 32 additions & 0 deletions docs/user/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -158,3 +158,35 @@ To connect OpenRouter:

> **Pausing.** Pausing the LLM runtime protects the local GPU, so it stops local models only.
> Hosted chat and hosted embeddings keep working while paused.

## A deployment with no local AI

Running **no local runtime at all** — hosted chat, hosted embeddings, no Ollama — is a
supported mode, not a broken install. It is spelled with a blank `OLLAMA_URL`: see
[`config` reference](../reference/config.md) for the variable and
[Kubernetes](../infrastructure/kubernetes.md) for the chart values that produce it.

What you see in the web UI when the core reports no local runtime:

- The **Models** page keeps only the hosted half. The catalog, the download tray, the local
model list, the default context window and the KV-cache card are *gone* — not greyed out —
and one line in their place says local AI is not configured on this deployment. Nothing there
could have done anything: every one of those cards is an Ollama control.
- The **Embedding model** picker offers no `Local (Ollama)` group, so you cannot pick a model
that has nothing to run on. Everything on offer is hosted, which means the whole notes,
knowledge and memory corpus goes to that provider — the trade-off described above is no
longer optional here, so make it deliberately.
- The **chat model picker** drops its `Local` heading and lists your hosted models. The *core
default* entry stays, because the core's default may itself be a hosted model.
- A module's **model slots** (Modules page) offer the core default and any saved hosted model;
a slot with neither says why it is empty rather than looking broken.

None of this is a warning, because nothing is wrong. A deployment that *does* expect a local
runtime and cannot reach it is a different state and still says so, in the words it always
used: "The local runtime is unreachable — is the ollama service up?" If you see that, something
is down; if you see "Local AI is not configured on this deployment", nothing is.

> **Set a hosted default.** With no local runtime, a stale `LLM_DEFAULT_MODEL` or
> `MEMORY_EMBED_MODEL` naming a local model (`llama3.2`, `nomic-embed-text`) cannot run. Star a
> hosted chat model under *Hosted models*, and pick a hosted embedding model under *Embedding
> model*, so both defaults point somewhere that answers.
4 changes: 2 additions & 2 deletions services/web/package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion services/web/package.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "epicurus-web",
"private": true,
"version": "0.150.0",
"version": "0.151.0",
"type": "module",
"description": "The epicurus web shell: chat, model manager, power toggle, and manifest-driven module UI.",
"scripts": {
Expand Down
5 changes: 5 additions & 0 deletions services/web/src/lib/api.ts
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ import {
FileText,
HoverCard,
LlmPrefs,
LocalRuntimeStatus,
LogEntry,
MaintenanceCurrentRun,
MaintenanceRunPage,
Expand Down Expand Up @@ -211,6 +212,10 @@ export const api = {
z.array(ModelInfo),
`/platform/v1/llm/models${withCapabilities ? "?capabilities=true" : ""}`,
),
// Whether this deployment runs a local runtime at all (#962) — `absent` (deliberately none),
// `unreachable` (one is configured and down), or `ok`. Its own endpoint rather than an
// envelope around the model list, which stays a bare array for its five web consumers.
localRuntime: () => request(LocalRuntimeStatus, "/platform/v1/llm/local-runtime"),
// The browsable model catalog the core parses from upstream on a schedule (#269).
catalog: () => request(CatalogResponse, "/platform/v1/llm/catalog"),
deleteModel: (name: string) =>
Expand Down
22 changes: 22 additions & 0 deletions services/web/src/lib/contracts.ts
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,28 @@ export const ModelInfo = z.object({
});
export type ModelInfo = z.infer<typeof ModelInfo>;

/**
* Whether this deployment runs a local model runtime at all (#962).
*
* Three states, because `absent` and `unreachable` are not the same thing and never were:
* `absent` is a **deliberate deployment mode** (`OLLAMA_URL=""` — hosted chat, hosted
* embeddings, no Ollama), so nothing is wrong and there is nothing for the operator to fix;
* `unreachable` is a runtime the deployment expects and cannot reach, which *is* an error and
* still looks like one. `ok` is the common case. The state is read from this endpoint and never
* inferred from a failing model list — since #962 `GET /llm/models` answers `[]` with a 200 in
* both of the unhappy states, so an error there no longer carries the information.
*/
export const LocalRuntimeState = z.enum(["absent", "unreachable", "ok"]);
export type LocalRuntimeState = z.infer<typeof LocalRuntimeState>;

export const LocalRuntimeStatus = z.object({
state: LocalRuntimeState,
// Whether `OLLAMA_URL` carries a value at all. Defaulted for tolerance: the state is what
// every surface renders from, and a core that reports one without the other is still usable.
url_configured: z.boolean().default(true),
});
export type LocalRuntimeStatus = z.infer<typeof LocalRuntimeStatus>;

/** Per-model tuning; null on a field means "inherit" the global / env default. */
export const ModelSettings = z.object({
context_window: z.number().nullable(),
Expand Down
63 changes: 63 additions & 0 deletions services/web/src/lib/useLocalRuntime.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
/**
* Does this deployment have a local model runtime? (#962)
*
* One query, one answer, read by every surface that used to assume Ollama existed. It exists
* because the shell had no way to tell three different things apart — a runtime that is *absent
* by design* (`OLLAMA_URL=""`: hosted chat, hosted embeddings, no Ollama), one that is
* configured and **unreachable**, and one that is fine — so it rendered the pull card, the
* catalog, the KV-cache card and a `Local (Ollama)` optgroup on a deployment where none of them
* could ever do anything, and called the whole situation "is the ollama service up?".
*
* Two rules this hook exists to keep:
*
* 1. **The state comes from `/llm/local-runtime`, never from a failing `/llm/models`.** Since
* #962 the model list answers `[]` with a 200 when the runtime is absent *or* unreachable,
* so an error there carries no information at all. Any code that goes back to inferring the
* state from a query failure silently stops working.
* 2. **An older core, or an unreachable one, reads as `ok`.** A core that predates the endpoint
* 404s; treating that as "absent" would hide the local half of the Models page from a
* deployment that has a perfectly good runtime. `ok` is exactly what every surface assumed
* before this change, so an unknown answer keeps today's behaviour.
*/
import { useQuery } from "@tanstack/react-query";

import { api } from "@/lib/api";
import type { LocalRuntimeState } from "@/lib/contracts";

export interface LocalRuntime {
/** `absent` | `unreachable` | `ok` — `ok` until the endpoint says otherwise. */
state: LocalRuntimeState;
/** No local runtime on this deployment, deliberately. The one flag most callers want. */
absent: boolean;
/** A runtime is configured and cannot be reached. An error, and it still looks like one. */
unreachable: boolean;
/** The endpoint has answered (or failed) at least once. Gate a *collapse* on this so a
* hosted-only deployment never flashes the local half of the page before hiding it again. */
settled: boolean;
}

export function useLocalRuntime(): LocalRuntime {
const status = useQuery({
queryKey: ["localRuntime"],
queryFn: () => api.localRuntime(),
// A deployment does not grow or lose a runtime between renders; one fetch per session is
// plenty, and a stale answer here is cheaper than polling a constant.
staleTime: 5 * 60_000,
// A 404 from an older core is a final answer, not a blip worth three attempts.
retry: false,
});
const state: LocalRuntimeState = status.data?.state ?? "ok";
return {
state,
absent: state === "absent",
unreachable: state === "unreachable",
// `isFetched` and not `!isPending`, because this gates whether the local half of the Models
// page is mounted at all, and those components subscribe to this very query. An errored
// query (an older core's 404) is permanently stale, so each newly-mounted subscriber
// triggers a refetch — which, with no data to fall back on, puts the query back into
// `pending` and unmounts them again: a page that oscillates forever. `isFetched` latches
// true after the first attempt, so the answer can change but the mounting decision can't
// feed back into it.
settled: status.isFetched,
};
}
68 changes: 52 additions & 16 deletions services/web/src/screens/ChatScreen.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,7 @@ import {
} from "@/lib/format";
import { SHARE_CACHE, SHARE_FILE_KEY, SHARE_FILE_NAME_HEADER, SHARE_META_KEY, type ShareMeta } from "@/lib/shareTarget";
import { SUGGESTION_VERB, suggestionTarget } from "@/lib/suggestions";
import { useLocalRuntime } from "@/lib/useLocalRuntime";
import {
fetchSessions,
useChat,
Expand Down Expand Up @@ -616,6 +617,10 @@ function ModelPicker() {
const sessionModel = sessions.data?.find((s) => s.id === sessionId)?.model ?? null;
const effectiveModel = sessionModel ?? model;
const models = useQuery({ queryKey: ["models"], queryFn: () => api.models(), enabled: open });
// No local runtime (#962) → no Local group at all: no heading, no "local runtime unreachable",
// nothing to pick. The core-default row survives on its own, because the core's default may
// perfectly well be a hosted model and this is the only way back to it.
const runtime = useLocalRuntime();
const providers = useQuery({ queryKey: ["providers"], queryFn: api.providers, enabled: open });
const llmPrefs = useQuery({ queryKey: ["llmPrefs"], queryFn: api.llmPrefs, enabled: open });
const saved = useQuery({ queryKey: ["savedModels"], queryFn: api.savedModels, enabled: open });
Expand Down Expand Up @@ -695,25 +700,34 @@ function ModelPicker() {
</span>
</p>
)}
<div>
<p className="mb-2 text-xs font-medium uppercase tracking-wide text-ink-faint">Local</p>
{runtime.absent ? (
<div className="flex flex-col gap-1">
<PickRow label={defaultLabel} active={effectiveModel === null} onPick={() => choose(null)} />
{visibleModels.map((m) => (
<PickRow
key={m.name}
label={m.name}
loaded={m.loaded}
size={m.size}
active={effectiveModel === m.name}
onPick={() => choose(m.name)}
/>
))}
{models.isError && (
<p className="text-xs text-warn">local runtime unreachable</p>
)}
</div>
</div>
) : (
<div>
<p className="mb-2 text-xs font-medium uppercase tracking-wide text-ink-faint">Local</p>
<div className="flex flex-col gap-1">
<PickRow label={defaultLabel} active={effectiveModel === null} onPick={() => choose(null)} />
{visibleModels.map((m) => (
<PickRow
key={m.name}
label={m.name}
loaded={m.loaded}
size={m.size}
active={effectiveModel === m.name}
onPick={() => choose(m.name)}
/>
))}
{/* Driven by the runtime's own state (#962) — the model list stopped erroring on
an unreachable runtime, so `models.isError` alone would never fire again. It
is kept beside it for an older core, which still 500s here. */}
{(runtime.unreachable || models.isError) && (
<p className="text-xs text-warn">local runtime unreachable</p>
)}
</div>
</div>
)}

<div>
<p className="mb-2 text-xs font-medium uppercase tracking-wide text-ink-faint">
Expand Down Expand Up @@ -817,6 +831,28 @@ function Welcome() {
const active = useDownloads((s) => s.active);
const queryClient = useQueryClient();
const suggestions = ["llama3.2", "qwen2.5:0.5b"];
// The first thing a new operator reads must not be advice they cannot take (#962): with no
// local runtime there is nothing to pull, and "pull a local one" would send them looking for a
// Pull button that this deployment deliberately doesn't have.
const runtime = useLocalRuntime();

if (runtime.absent) {
return (
<EmptyState quote={dayQuote()}>
<Card className="mt-2 w-full max-w-sm text-left">
<h3 className="font-serif text-base text-ink">Welcome to the garden</h3>
<p className="mt-1 text-sm leading-relaxed text-ink-dim">
No model is configured yet. This deployment runs no local AI, so add a hosted
provider key under{" "}
<Link to="/models" className="text-accent-strong underline">
Models
</Link>{" "}
and pick a model there.
</p>
</Card>
</EmptyState>
);
}

return (
<EmptyState quote={dayQuote()}>
Expand Down
Loading
Loading