Skip to content

feat: OpenAI-compatible inference backend alongside Ollama - #7

Merged
wikithoughts merged 1 commit into
mainfrom
feat/openai-compatible-inference-backend
Aug 30, 2026
Merged

wikithoughts merged 1 commit into
mainfrom
feat/openai-compatible-inference-backend

Conversation

@wikithoughts

Copy link
Copy Markdown
Owner

What

Adds PW_INFERENCE_BACKEND=ollama|openai + PW_INFERENCE_API_BASE to council/ollama.py, the single shared HTTP client used by worker, researcher, judge, batch, and rerank. Since every one of those callers already goes through this one module's base() / generate() / smallest_chat_model(), this gives all five a new backend option with zero changes to those callers.

  • Default (ollama, or unset): behavior unchanged — /api/generate + /api/tags, PW_OLLAMA_BASE.
  • PW_INFERENCE_BACKEND=openai: speaks the OpenAI-compatible /v1/chat/completions + /v1/models shape that both llama.cpp-server and LM Studio expose out of the box (2026), resolving the endpoint from PW_INFERENCE_API_BASE (no hardcoded default — there's no sane guess for an arbitrary local port the way localhost:11434 is for Ollama).

generate()'s public signature and return shape (tuple[str, int] of (text, tokens)) are unchanged; internally it now dispatches to _generate_ollama/_generate_openai. The openai path mirrors the existing Ollama path's error-tolerance contract exactly (0 tokens / empty text on a missing field, never a crash).

Out of scope (intentional)

Other call sites that talk to Ollama directly (council/local.py, council/library.py, council/doctor.py, council/net/agent.py, council/net/baseline.py, council/operator.py) are not routed through this dispatch and remain Ollama-only — documented as a known limitation in the module's docstring, not an oversight. D1 (no token), D4 (no proxied traffic), and D18 (no browser automation) are unaffected — this is purely a local-inference-runtime choice.

Config

Registered PW_INFERENCE_BACKEND and PW_INFERENCE_API_BASE in council/config.py's KNOWN registry so they're discoverable via pw config list.

Tests

New tests/test_ollama.py (module had zero dedicated tests before this PR): base() resolution on both backends, generate()'s request shape on both backends (mocked requests.post, asserted on URL/body), resolve_timeout() precedence, and the two new KNOWN keys.

Verification

  • python -m py_compile council/ollama.py council/config.py tests/test_ollama.py — clean
  • ruff check . — all checks passed
  • pytest tests/ -q — 521 passed, 1 skipped, 0 failed (full suite, including everything touching council/ollama.py / council.config.KNOWN)
  • No new third-party dependency — only stdlib os and the already-core requests package, so the core-install-no-extras CI smoke job is unaffected.

🤖 Generated with Claude Code

…Studio)

Adds PW_INFERENCE_BACKEND=ollama|openai + PW_INFERENCE_API_BASE to
council/ollama.py, the single shared HTTP client used by worker,
researcher, judge, batch, and rerank. On the default (ollama) backend
behavior is unchanged. On openai, base()/generate()/smallest_chat_model()
speak the OpenAI-compatible /v1/chat/completions + /v1/models shape that
llama.cpp-server and LM Studio expose, giving every council/ollama.py
caller a new backend option with zero changes to those callers.

Other direct-to-Ollama call sites (council/local.py, council/library.py,
council/doctor.py, council/net/agent.py, council/net/baseline.py,
council/operator.py) are intentionally out of scope and remain
Ollama-only, per the task boundary.

Adds tests/test_ollama.py (module had zero dedicated tests) covering
base() resolution on both backends, generate()'s request shape on both
backends, resolve_timeout() precedence, and KNOWN registry entries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@wikithoughts
wikithoughts merged commit e3cdb41 into main Aug 30, 2026
6 of 9 checks passed
@wikithoughts
wikithoughts deleted the feat/openai-compatible-inference-backend branch August 31, 2026 10:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant