Local multi-backend proxy that speaks the Anthropic Messages API so Claude Code (CLI + VSCodium/VS Code) can run on:
- OAuth providers — xAI Grok, Kimi Code, optional chat.qwen.ai, OpenAI (ChatGPT/Codex) (
spock login <provider>) - API Key backends — Qwen Cloud (qwencloud.com), Ollama / llama-server / OpenRouter / any OpenAI-compatible API
- Responses backends — the OpenAI ChatGPT / Codex subscription (see ChatGPT subscription) — no separate API billing
Claude Code always points at Spock (http://127.0.0.1:8048). Spock maps Haiku / Sonnet / Opus / Fable (and any model id) to different backends via profiles — without changing Claude settings when you switch vendors.
| Piece | Role |
|---|---|
| Spock.app | macOS menu bar app — proxy, Settings, Chat, profile switch, status icon |
spock CLI |
serve, login, logout, chat, status, reload (all platforms) |
| Config | ~/.config/spock/config.toml |
| OAuth tokens | ~/.config/spock/oauth-<provider>.json (imports legacy grok-test / kimi paths once) |
Spock is protocol-compatible with Claude Code’s Anthropic Messages client. Versions below were verified working with this Spock release (normal chat, tools, streaming, profile routing). Future Claude Code releases can change betas, tool schemas, or streaming and break a proxy — pin or re-test when upgrading.
| Spock | Claude Code CLI | Claude Code IDE extension | Host (example) | Status | Notes |
|---|---|---|---|---|---|
| 0.3.0 | 2.1.233 | 2.1.233 | VSCodium / VS Code | OK | Vision policy + KV sessions + catalog UI (2026-08-24) |
| 0.2.0 | 2.1.207 | 2.1.207 (cc_version=2.1.207.6dd) |
VSCodium 1.128.0 (also VS Code) | OK | Microcompact + mid-SSE errors + log-file + webview text server tools (2026-07-13) |
| 0.2.0 | 2.1.206 | 2.1.206 (cc_version=2.1.206.87c) |
VSCodium 1.128.0 (also VS Code) | OK | Server-tool emulation + presets + Auto Mode reasoning_effort fix (2026-07-13) |
| 0.1.0 | 2.1.206 | 2.1.206 | VSCodium 1.128.0 | OK | Baseline Rust multi-backend release (2026-07-11) |
How to check your versions
# Spock
curl -s http://127.0.0.1:8048/health | jq -r .version
# Claude Code CLI
claude --version
# VSCodium / VS Code extension (folder name includes version)
ls -d ~/.vscode-oss/extensions/anthropic.claude-code-* 2>/dev/null
ls -d ~/.vscode/extensions/anthropic.claude-code-* 2>/dev/nullIf something breaks after a Claude Code upgrade: note both Spock and Claude Code versions, re-run the smoke test below, and check Troubleshooting. Prefer matching CLI and extension versions.
Known limits on the verified stack (not full breakage)
- Server-tool emulation is opt-in via
[advisor]/[web_search]in config (defaults off). Without them,advisor_20260301/web_search_*schemas are stripped for OpenAI-compat upstreams. - Advisor runs as a nested review, then comes back as a text block (
Advisor review:…). VSCodium/VS Code 2.1.226 webview has no renderer forserver_tool_use/advisor_tool_resultand prints those names as chat lines. Never returned as a clienttool_use(No such tool available: advisor). Pin[advisor].modelto a different route than the executor or the reviewer just continues the agent voice. History that still has the protocol blocks is flattened on the OpenAI-compat path. - OpenAI Responses API flag exists but is not implemented — leave
use_responses_api = false(Chat Completions). - Client-side microcompact runs when Claude Code sends
context_managementedits; Anthropic passthrough leaves them for real Anthropic.
- From Releases, download
Spock-VERSION-darwin-arm64.zip(Apple Silicon) or…-darwin-x64.zip - Unzip → drag Spock.app to Applications
- First launch: right-click → Open if Gatekeeper warns (ad-hoc signed until Developer ID is set)
- Menu bar icon appears (no Dock icon — menu bar agent)
- Settings… — backends + profiles (xAI / Ollama / OpenAI-compat / …). Optional: Login xAI… or paste an xAI API key if you use Grok.
Build from source:
./packaging/macos/build-app.sh # → dist/Spock.app
open dist/Spock.app# From a release tarball
tar -xzf spock-VERSION-darwin-arm64.tar.gz
sudo mv spock-VERSION-darwin-arm64/spock /usr/local/bin/
# From source
cargo build --release
# binary: target/release/spock# Option A — macOS app (starts proxy automatically)
open dist/Spock.app # or Spock from Applications
# Option B — CLI
spock login xai # Grok subscription (or: kimi / qwen)
spock serve # http://127.0.0.1:8048Smoke test:
curl -s http://127.0.0.1:8048/health | jq
curl http://127.0.0.1:8048/v1/messages \
-H 'content-type: application/json' \
-d '{"model":"grok-4.5","max_tokens":256,"messages":[{"role":"user","content":"hello"}]}' | jqMenu bar only (LSUIElement / activation policy .accessory):
| Menu | Action |
|---|---|
| Status line | Profile · port · proxy state |
| Chat… | Native chat window against the proxy |
| Settings… | Full config UI (backends, profiles, routes, API keys) |
| Profile | Switch active profile live |
| Reload config | Re-read config.toml |
| Login xAI… / Logout xAI | OAuth device flow |
| Quit Spock | Stop proxy (if app started it) and exit |
Status icon color
| Color | Meaning |
|---|---|
| Green | Proxy healthy |
| Orange | Starting |
| Gray | Stopped |
| Red | Error |
Closing Settings/Chat does not quit the app (no Dock icon stuck around). Only Quit Spock exits.
Settings highlights
- Active profile (persists immediately on change)
- Backends: OAuth / API Key / Anthropic rows, base URL, optional API key, per-backend Text-only upstream tick
- Fetch models — pulls model ids from any OpenAI-compatible backend
- Profiles & routes —
backend:modelper role, with dropdowns from fetched models - Catalog — curated
GET /v1/modelsshortlist for external pickers (Grok Build), with context + effort columns - Server tools —
[advisor]/[web_search]emulation toggles - Vision sidecar —
mode = strip|describe+ endpoint for text-only backends - Save & Apply — writes TOML and hot-reloads the proxy
Providers are registered in Spock (spock providers). Each OAuth backend is:
[backends.xai]
type = "oauth"
provider = "xai"
[backends.kimi]
type = "oauth"
provider = "kimi"
Auth priority per OAuth provider (first wins):
- API key on that backend (Settings / config) — escape hatch
- Env token (
XAI_TOKEN,KIMI_TOKEN, …) - OAuth device login
Tokens:
spock login xai spock login kimi # menu: Login ▸ xAI / Kimi Code~/.config/spock/oauth-<provider>.json(0600).
spock logout xai
spock logout kimi
spock logout --allqwencloud.com has two paid plans with different hosts and keys. Neither is chat.qwen.ai OAuth.
| Plan | Key | OpenAI-compatible base | Models (examples) |
|---|---|---|---|
| Coding Plan | sk-sp-… |
https://coding-intl.dashscope.aliyuncs.com/v1 |
Fixed allowlist: qwen3.7-plus, qwen3.6-plus, qwen3-coder-plus, glm-5, kimi-k2.5, … — no qwen3.8-max-preview |
| Token Plan | separate key (create on Token Plan page) | https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 |
Credits catalog incl. qwen3.8-max-preview, qwen3.7-max, … |
# Coding Plan (sk-sp-…)
[backends.qwen]
type = "api_key"
base_url = "https://coding-intl.dashscope.aliyuncs.com/v1"
api_key_env = "DASHSCOPE_API_KEY"
# Token Plan (for qwen3.8-max-preview) — different key + host:
# [backends.qwen-token]
# type = "api_key"
# base_url = "https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1"
# api_key = "…" # from Token Plan page — NOT your Coding Plan sk-sp key- Create the matching key: API Keys / Token Plan
- Paste into Settings or env; do not mix Coding Plan keys with the Token Plan host (or vice versa)
- Route e.g. Coding:
fable = "qwen:qwen3.7-plus"· Token:fable = "qwen-token:qwen3.8-max-preview" - Reload Spock / Save & Apply
/models only lists what that plan’s host exposes. Docs can advertise qwen3.8-max-preview while a Coding Plan key’s catalog stays at 10 fixed models — that is upstream plan gating, not Spock.
Optional: spock login qwen is the chat.qwen.ai OAuth path (qwen-code). Different product again.
Your ChatGPT subscription is not the OpenAI platform API — the platform rejects the chatgpt token, and subscription-gated models (e.g. gpt-5.5) are only reachable through the Codex consumer backend (POST https://chatgpt.com/backend-api/codex/responses), which speaks the OpenAI Responses API and is stream-only. Spock hides that wire: openai is a plain OAuth provider — you configure it exactly like xAI/Kimi.
[backends.chatgpt]
type = "oauth"
provider = "openai"
base_url = "https://chatgpt.com/backend-api"No Codex app is required. Login is a real OAuth authorization-code PKCE flow — spock login openai (or menu-bar Login ▸ OpenAI) opens a browser, you sign in, and Spock exchanges the code for tokens and stores them in its own ~/.config/spock/oauth-openai.json. The browser passes any Cloudflare challenge, so the auth.openai.com device-flow-blocking does not apply; Spock keeps its own token alive via auth.openai.com/oauth/token. If you previously used Codex, the import from ~/.codex/auth.json is used only as a convenience; once logged in via Spock it self-sustains.
Optionally, for an OpenAI API key (separate billing, official gateway) use the type = "responses" kind with base_url = "https://api.openai.com/v1" + path = "responses" + api_key_env = "OPENAI_API_KEY".
spock login openai(one-time; browser opens) — or leave it if you already have a Codex login- Route e.g.
fable = "chatgpt:gpt-5.5"·opus = "chatgpt:gpt-5.5" - Reload Spock / Save & Apply. Model ids:
gpt-5.5,gpt-5.4,gpt-5.2, … (effortlow…xhighviareasoning_effort).
[backends.ollama]
type = "api_key"
base_url = "http://127.0.0.1:11434/v1"Platform Moonshot metered keys use api.moonshot.ai as type = "api_key" — that is separate from Kimi Code OAuth (api.kimi.com/coding/v1).
Claude Code keeps a single base URL:
ANTHROPIC_BASE_URL=http://127.0.0.1:8048
Spock resolves each request’s model id using the active profile:
exact id override
→ role (haiku / sonnet / opus / fable) if the id contains that word
→ profile default
→ backend:upstream_model
There is no automatic grok* → xAI shortcut. A client id like grok-4.5[1m] uses the profile default (or a role row if you map it). Context suffixes ([1m], [500k], …) are stripped before the upstream call; Claude Code still sees the original id in responses.
Example (config.example.toml):
[server]
profile = "hybrid"
[backends.xai]
type = "oauth"
provider = "xai"
[backends.ollama]
type = "api_key"
base_url = "http://127.0.0.1:11434/v1"
[profiles.hybrid]
default = "ollama:glm-5.2:cloud"
haiku = "ollama:kimi-k2.7-code:cloud"
sonnet = "ollama:kimi-k2.7-code:cloud"
opus = "xai:grok-4.5"
fable = "ollama:glm-5.2:cloud"| Claude Code sends (examples) | Hybrid row used |
|---|---|
id containing haiku |
haiku |
id containing sonnet |
sonnet |
id containing opus |
opus |
id containing fable |
fable |
grok-4.5[1m], other unmatched ids |
default |
Tip: Different backends have different “personalities.” Grok often follows Claude Code’s “you are Claude” system prompt; other models may answer as themselves (or in another language). For a single-brain session, point default + fable + main roles at the same backend:model.
Many public OpenAI-compatible gateways already work as type = "api_key" with the right base_url + api_key (or api_key_env). Examples (see config.example.toml):
| Backend name | base_url | Key |
|---|---|---|
| OpenRouter | https://openrouter.ai/api/v1 |
OPENROUTER_API_KEY or api_key |
| OpenAI | https://api.openai.com/v1 |
OPENAI_API_KEY |
| DeepSeek | https://api.deepseek.com |
DEEPSEEK_API_KEY |
| Groq | https://api.groq.com/openai/v1 |
GROQ_API_KEY |
| LM Studio | http://127.0.0.1:1234/v1 |
optional |
Optional on openai backends:
api_key_env = "OPENROUTER_API_KEY"
[backends.openrouter.extra_headers]
HTTP-Referer = "https://github.com/satindergrewal/Spock"
X-Title = "Spock"Then route e.g. default = "openrouter:anthropic/claude-sonnet-4". No Claude Code changes required.
Switch profiles live: app menu Profile, Settings active profile, or edit TOML + Reload / spock reload.
Config path: ~/.config/spock/config.toml (created on first run). Full sample: config.example.toml.
Text-only upstreams (vLLM DSV4-Flash, GLM-5.3, …) hard-400 on image content, and clients re-send the full transcript — one screenshot poisons every later request in the session. Flag the backend and Spock rewrites images before the request leaves:
[backends.lan]
type = "api_key"
base_url = "http://10.0.0.5:8080/v1"
text_only = true
[vision]
mode = "describe" # or "strip" (default)
sidecar_base_url = "http://10.0.0.5:8090/v1"
sidecar_model = "your-vl-model"
timeout_secs = 30strip— screenshots become[image omitted: this backend is text-only].describe— each screenshot is captioned by a small VL sidecar (any OpenAI-compatible vision endpoint, e.g. llama-server + mmproj) and the caption is inlined as text. Any sidecar failure degrades to strip; a request never dies here.- Captions are cached in memory only (sha256 of image + prompt), so re-sent history is free. No per-request image limit — one failed sidecar call strips the rest of that request; a healthy sidecar captions every image.
- Retroactively un-sticks poisoned sessions: the next request goes out clean, no new chat needed.
- Settings UI: per-backend Text-only upstream tick + Vision section. Unflagged / vision backends pass images through untouched.
kv_sessions = true on an api_key backend parks a named master context on llama-server's native /completion + /fork + /close_session routes instead of re-prefilling every turn. Missing routes or unknown sessions error the request — never a silent fall-through to /v1/chat/completions.
Let Spock own models via the profile (recommended for hybrid):
"claudeCode.environmentVariables": [
{ "name": "ANTHROPIC_BASE_URL", "value": "http://127.0.0.1:8048" },
{ "name": "ANTHROPIC_API_KEY", "value": "xai" },
{ "name": "ANTHROPIC_AUTH_TOKEN", "value": "xai" },
{ "name": "CLAUDE_CODE_AUTO_COMPACT_WINDOW", "value": "500000" },
{ "name": "CLAUDE_AUTOCOMPACT_PCT_OVERRIDE", "value": "90" },
{ "name": "API_TIMEOUT_MS", "value": "3000000" }
]Optional model env overrides (only if you want Claude Code to send fixed ids):
{ "name": "ANTHROPIC_MODEL", "value": "claude-fable-5[1m]" },
{ "name": "ANTHROPIC_DEFAULT_FABLE_MODEL", "value": "claude-fable-5[1m]" },
{ "name": "ANTHROPIC_DEFAULT_OPUS_MODEL", "value": "claude-opus-4-8" },
{ "name": "ANTHROPIC_DEFAULT_SONNET_MODEL", "value": "claude-sonnet-5" },
{ "name": "ANTHROPIC_DEFAULT_HAIKU_MODEL", "value": "claude-haiku-4-5" }Those ids are role tags for Spock routing — not necessarily the real upstream model names.
Start a new Claude Code session after changing env.
Use a Claude Code version from the compatible table when possible.
./claude-grok.shStarts spock serve if needed, sets compact window env, launches claude. Override model with GROK_MODEL_ID / ANTHROPIC_MODEL if you want.
- Pair a
[1m]-flavoured client id withCLAUDE_CODE_AUTO_COMPACT_WINDOW=500000so compact math matches Grok’s ~500k window. - Spock strips
[1m]/[500k]before calling upstream. - Optional:
CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=90.
xAI reasoning models reject OpenAI stop. Spock drops stop / presence / frequency penalties on xAI reasoning routes and keeps them for Ollama.
GET /v1/models always includes Claude aliases so Auto Mode does not treat the classifier model as missing.
spock serve [--port N] Headless proxy on 127.0.0.1 (default 8048)
spock app Open Spock.app (macOS)
spock login <provider> [--no-open] OAuth device login (xai, kimi, qwen, …)
spock logout <provider> | --all
spock providers [--json]
spock chat [prompt] [-m model]
spock status Profile, backends, auth source, proxy health
spock reload Re-read config.toml
spock -V | --version
spock help
| Variable | Default | Purpose |
|---|---|---|
PORT |
8048 |
Listen port |
GROK_MODEL |
grok-4.5 |
Legacy default / alias helper |
GROK_SMALL_MODEL |
= GROK_MODEL |
Haiku-style alias helper |
XAI_TOKEN |
— | xAI API key (skips OAuth if set) |
XAI_API_BASE |
https://api.x.ai/v1 |
Upstream override |
| Endpoint | Notes |
|---|---|
POST /v1/messages |
Anthropic Messages (stream + tools + thinking) |
POST /v1/messages/count_tokens |
Rough estimate (chars/4) |
POST /v1/chat/completions |
OpenAI-style (raw stream passthrough) |
POST /v1/responses |
Search-only (grok-build web_search). Runs [web_search]; not a general Responses proxy |
GET /v1/models |
Curated catalog when configured, else merged backend lists + Claude aliases |
GET /v1/models/{id} |
Never 404s for aliases |
GET /v1/language-models* |
xAI extended list when available |
GET /health |
Status, profile, backends, version |
| Endpoint | Notes |
|---|---|
GET /spock/v1/status |
Profile, auth source, paths |
GET /spock/v1/config |
Settings document |
PUT /spock/v1/config |
Save & apply full config |
POST /spock/v1/profile |
{"profile":"hybrid"} |
POST /spock/v1/reload |
Reload from disk |
POST /spock/v1/logout |
Clear OAuth file |
GET /spock/v1/backends/{name}/models |
Discover models (Ollama / xAI) |
# Tests + headless binary
cargo test
cargo build --release
# macOS app (Rust proxy + SwiftUI shell)
./packaging/macos/build-app.sh
# → dist/Spock.appTag vX.Y.Z must match Cargo.toml version. Pushing the tag builds:
- CLI archives: darwin-arm64/x64, linux-x64/arm64, windows-x64
- macOS App zips:
Spock-VERSION-darwin-arm64.zip/…-x64.zip checksums.txt+ release notes
Workflows: .github/workflows/ci.yml, .github/workflows/release.yml.
- Binds 127.0.0.1 only — anyone who can reach the port can use your backends
- OAuth token file
0600on Unix - Do not commit real API keys, LAN IPs, or tokens to public files
- Admin API is unauthenticated by design (loopback only)
| Symptom | Fix |
|---|---|
| Need proxy logs | App: tail -f ~/Library/Logs/Spock/spock.log · CLI: spock serve --log-file /tmp/spock.log |
| Upstream quota / 401 in IDE | Settings error banner / spock status → last_err; fix backend key/credits (not Anthropic login) |
401 / SuperGrok / usage on xAI |
Check API key, XAI_TOKEN, or spock logout && spock login; quota on the xAI side |
| Claude opens “log in to Anthropic” | Often a misread upstream error — check Spock logs; use a valid key/OAuth; prefer latest Spock (upstream 401 → 502 with clear message) |
| Active profile snaps back in Settings | Use latest app (profile switch persists immediately); Save & Apply |
| Hybrid hits wrong model | Check active profile rows; fable/default often drive main chat; empty role fields fall through to default |
Ollama *:cloud fails |
Sign in / enable that model in Ollama cloud |
LAN Fetch models: No route to host (os error 65) while Terminal curl works |
macOS Local Network privacy. Rebuild app (./packaging/macos/build-app.sh includes NSLocalNetworkUsageDescription), Quit Spock, reopen. System Settings → Privacy & Security → Local Network → enable Spock. Or run spock serve from Terminal (inherits Terminal’s LAN access). |
type = "oauth" parse error from dist/Spock.app |
App binary is old — rebuild with ./packaging/macos/build-app.sh; do not hot-copy spock-proxy into a signed .app |
| Address already in use | lsof -nP -iTCP:8048 — quit old Spock or Python proxy |
| Gatekeeper blocks app | Right-click → Open |
| Auto Mode “temporarily unavailable” | Upgrade Spock; ensure stop is dropped for xAI reasoning; aliases on /v1/models |
| Compacts too early | CLAUDE_CODE_AUTO_COMPACT_WINDOW=500000 + optional [1m] client id |
| Broke after Claude Code update | Compare your CLI/extension versions to the compatible table; pin last good extension while Spock is updated |
| xAI / backend quota or 401 in VSCodium | Prefer latest Spock — upstream 401/402/403/429 → 502 with a loud Spock upstream … message (not Anthropic login). Check spock status / xai_auth, credits on console.x.ai, or switch profile |
tool_choice set but no tools |
Fixed when tools are stripped (server tools); upgrade Spock |
| Auto Mode “claude-opus-… unavailable” | Often dead opus route or (older Spock) reasoning_effort: none. Point opus at a live backend; use Spock with the thinking-disabled fix |
upstream 400: … is not a multimodal model |
Text-only backend received a screenshot. Set text_only = true on that backend (see Vision); the poisoned session heals on the next request |
Claude Code ECONNRESET / Grok Build reqwest error stream: error sending request |
Darwin accept() used to inherit O_NONBLOCK from the listen socket. Rebuild (./packaging/macos/build-app.sh) and restart Spock. If it still happens: spock.log should now show route … [openai+stream] and spock openai stream: … instead of a silent drop. Direct LAN vs SSH tunnel is a separate hop. |
Proxy logs each request with the resolved route, e.g.:
POST /v1/messages
route claude-fable-5[1m] → ollama:glm-5.2:cloud (openai)
- Rust proxy: minimal deps (
ureq,serde,toml,sha2), threadedTcpListener, Anthropic ↔ OpenAI translation - SwiftUI macOS shell: menu bar + Settings + Chat; talks to proxy admin API
- Profiles hot-reload without restarting Claude Code
Python implementation has been removed; this repo is Rust + Swift only.
MIT
Thanks for using Spock — run Claude Code on the models you choose, locally.
