Repository navigation
Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) - #136
Merged
Merged
Conversation
…de of #53 (D-058) Port upstream's chat.py for the in-process generator (MlxGenerator), against a stub, while the model's generation is ported in parallel (#50 to #52). OpenJevCore declares TextGenerator, whose generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) is MlxRuntime.generate, with the prompt, stop encoding, thought-channel markers and prompt limit the route reads; the server adds POST /v1/chat/completions when SystemOneService.textGenerator is set, which DecisionEngine gives for a backend that conforms. The route checks, normalizes, bounds capacity (529, retry-after 2), refuses an over-long prompt before the answer starts, answers whole replies with usage and JSON mode's extraction, and streams OpenAI's chunks through a queue of 64 pieces; a client that goes away, or a reader that falls 64 pieces behind, stops the generation at its next block, and a reply that ends early is always a prefix. SwiftTransformersTokenizer renders a chat request's messages as upstream's prompt_ids does, the scaffold after, on a 16 MB thread, with jinja2's sequence test and Python's str() in trim. Tools/fixtures/chat_tables.py records upstream's normalize, extract_json, prompt_ids and 69 route exchanges (Fixtures/chat-completions); the port matches 61 byte for byte, event streams included, and the other 8 are the departures D-058 records. The tests carry upstream's names; the ones that need the model stay disabled, naming #51. Refs #53.
Item 6 adds how a graceful and a timed-out shutdown treat a stream in flight, which the route now handles and tests; item 11 says which disabled cases have bodies and reflows two paragraphs. Refs #53.
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unbounded rendering threads, Unicode-equivalent key loss, and an unlogged backend failure path need correction.
Review effort: Balanced
Findings: 1
Open (3)
What changed in this PR
Adds the core and HTTP infrastructure for OpenAI-compatible chat completions, ready for DiffusionGemma generation wiring.
Changes:
- Introduces request normalization, generation capacity, streaming, JSON extraction, and response/error contracts.
- Adds conditional server routing, cancellation handling, tokenizer prompt rendering, and fixture-backed parity tests.
- Documents chat configuration, compatibility, architecture, and deployment behavior.
| File | Description |
|---|---|
Sources/OpenJevCore/Generation/TextGenerator.swift |
Defines generation protocol and result contract. |
Sources/OpenJevCore/Generation/ChatCompletionRequest.swift |
Validates and normalizes chat requests. |
Sources/OpenJevCore/Generation/ChatCompletions.swift |
Implements preparation, capacity, and completion. |
Sources/OpenJevCore/Generation/ChatCompletionStream.swift |
Implements bounded SSE streaming. |
Sources/OpenJevCore/Generation/ChatCompletionResponse.swift |
Defines replies, usage, and SSE events. |
Sources/OpenJevCore/Generation/ChatCompletionError.swift |
Defines OpenAI-shaped errors. |
Sources/OpenJevCore/Generation/ExtractJSON.swift |
Extracts JSON-mode responses. |
Sources/OpenJevCore/Engine/SystemOneService.swift |
Exposes optional text generation. |
Sources/OpenJevCore/JSON/JSONValue+Python.swift |
Adds Python truthiness and representations. |
Sources/OpenJevCore/JSON/PythonJSONWriter.swift |
Adds non-finite-number output support. |
Sources/OpenJevCore/Schema/TextOf.swift |
Adds Python-compatible right trimming. |
Sources/OpenJevServer/ChatCompletionsRoute.swift |
Serves whole and streamed completions. |
Sources/OpenJevServer/OpenJevApplication.swift |
Conditionally registers the chat route. |
Sources/OpenJevServer/WireResponses.swift |
Renders chat errors. |
Sources/OpenJevServer/RefusalLog.swift |
Logs generation failures. |
Sources/OpenJevServer/Documentation.docc/Configuration.md |
Documents generation settings. |
Sources/OpenJevDiffusionGemma/Tokenization/SwiftTransformersTokenizer.swift |
Builds generation prompts and Jinja parity behavior. |
Sources/OpenJevTestSupport/StubTextGenerator.swift |
Supplies stub generation runtimes. |
Sources/OpenJevTestSupport/StubGeneratingBackend.swift |
Combines stub reads and generation. |
Sources/OpenJevTestSupport/ChatFixtures.swift |
Loads recorded chat fixtures. |
Tests/OpenJevCoreTests/Generation/ChatCompletionsTests.swift |
Tests capacity, streaming, and replies. |
Tests/OpenJevCoreTests/Generation/ChatFixtureTests.swift |
Tests normalization and JSON extraction parity. |
Tests/OpenJevServerTests/Support/ChatHarness.swift |
Provides chat route test infrastructure. |
Tests/OpenJevServerTests/ChatCompletionsRouteTests.swift |
Tests the HTTP contract. |
Tests/OpenJevServerTests/ChatRouteFixtureTests.swift |
Replays recorded exchanges. |
Tests/OpenJevServerTests/ChatLiveServerTests.swift |
Tests live cancellation and shutdown. |
Tests/OpenJevDiffusionGemmaTests/Tokenization/ChatTemplateTrimTests.swift |
Tests Python/Jinja semantics. |
Tests/OpenJevDiffusionGemmaTests/Tokenization/ChatCompletionPromptTests.swift |
Tests prompt parity and nesting. |
Tests/OpenJevDiffusionGemmaTests/Runtime/UpstreamReadCaseTests.swift |
Adds disabled model-backed chat cases. |
Tests/OpenJevDiffusionGemmaTests/Runtime/RuntimeTests.swift |
Verifies generation remains unwired. |
Tests/OpenJevLiveTests/LiveTests.swift |
Updates live-test prerequisites. |
Tools/fixtures/chat_tables.py |
Generates upstream chat fixtures. |
Tools/README.md |
Documents the fixture generator. |
Makefile |
Runs chat fixture generation. |
Fixtures/README.md |
Catalogues chat fixtures. |
Fixtures/chat-completions/README.md |
Documents fixture contents. |
Fixtures/chat-completions/prompts.json |
Records prompt parity cases. |
Fixtures/chat-completions/normalize.json |
Records normalization cases. |
Fixtures/chat-completions/extract_json.json |
Records JSON extraction cases. |
Fixtures/chat-completions/routes.json |
Records HTTP exchanges. |
README.md |
Updates chat availability status. |
THIRD_PARTY.md |
Records CPython-derived scanner behavior. |
docs/05-architecture.md |
Describes chat architecture. |
docs/06-decisions.md |
Records D-058 decisions. |
docs/09-conformance-and-testing.md |
Documents parity coverage. |
docs/compatibility.md |
Updates compatibility claims. |
docs/deployment.md |
Documents chat operation and settings. |
docs/development.md |
Updates fixture commands. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
…ja2 does The first review of PR #136 (Copilot and an independent review): - ChatCompletions.prepare counts a request against the capacity bound at the check, before its prompt renders, and gives the place back when it is refused afterwards. Upstream counts after rendering, but renders on its event loop; here each prompt renders on a 16 MB thread of its own, which a burst could otherwise start without limit. PreparedChatCompletion carries the place and is answered once. The bound saturates at Int.max instead of trapping. - The route logs the 503 prepare gives for a stop string the generator cannot encode, as it logs a generation's. - The chat template's dictsort is jinja2's: keys ordered by Python's str.lower(), final sigma included, then code point, where swift-jinja's compares them with Foundation's case-insensitive compare. Two prompts.json rows record upstream's rendering of such keys. - An object whose keys differ only in Unicode normalization is a 400: swift-jinja keys objects by Swift's String and would keep one of them. - The live disconnect test waits for the next request's 200 rather than the stub's count, which drops before the stream gives its place back. Refs #53.
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
…Generator (#53, D-059) The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10.
This was referenced Oct 6, 2026
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
…l then" (#140) PR #136 added `extension ChatCompletionsConfiguration` with an `init(_:)` in Sources/OpenJevServer/ChatCompletionsRoute.swift, next to the extensions of EngineConfiguration and EncoderEngineConfiguration in BackendProvider.swift. Those are the only three extensions in Sources/OpenJevServer, and OpenJevCore declares all three types. docs/development.md's Links bullet and the header of Tools/docs/build-site.sh still said two; both say three now, and development.md names the three `init(_:)` the server adds. GitHub Pages publishes from GitHub Actions: the Pages API reports build_type workflow, and the Documentation run on main for d25223a deployed the site. The header of .github/workflows/docs.yml now says the deploy job is skipped "otherwise" instead of "until then", as PR #139 words docs/development.md, whose own copy of that line PR #139 changes. docs/06-decisions.md keeps its wording: a decision records what was true when it was written.
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
Every public symbol declared in the five library modules now has a doc comment, as the Coverage bullet of docs/development.md says. Comments only: no code, signature or access level changes. - VisionTower.swift (#129): the @ModuleInfo and @ParameterInfo properties and the callAsFunction methods of VisionAttention, VisionMLP, VisionBlock, VisionPatchEmbedder, VisionModel, VisionEncoder and MultimodalEmbedder (29), and the callAsFunction of ClippableLinear and VisionRMSNorm. A property says what it holds and its checkpoint key, as the text model's do; a callAsFunction says what it computes, citing mlx-vlm's __call__, with the file's shapes. - VisionError.description (#121): "The message.", as the other errors' descriptions say. - extension ChatCompletionsConfiguration in OpenJevServer (#136): a /// line on the extension, as BackendProvider.swift's two extensions of core types have.
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
…er, the block loop, think and the chat model (#50, #51, #52, #53, D-059) (#137) * Generate text and think on the MLX backend: the diffusion sampler, the block loop and the think option (#50, #51, #52, D-058) Port mlx-vlm 0.6.15's sampling functions (DiffusionSampler), one block of stream_diffusion_generate as upstream's MlxRuntime.generate runs it (greedy, the default confidence-threshold sampler, the linear temperature schedule, self-conditioning, stable-and-confident stopping), and diffusion_update_cache on copies of the layer caches. DiffusionGemmaRuntime gains generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) with canvas sizing, EOS and stop ids, the SentencePiece streaming detokenizer, emit and task cancellation, and think(prompt:budget:stopIDs:) on it, so capabilities.think is on. Each reply's canvases come from a RandomState seeded with Configuration.generationSeed (0), which upstream leaves unseeded. Tools/fixtures/generation_oracle.py records the oracle in Fixtures/generation: sampler vectors, seven seeded greedy replies block by block, and two think requests. On the oracle's Metal library every reply and thought matches token for token with upstream's billing; natively the short replies match whole and the long story for its first 8 tokens. Upstream's three think cases and the live suite's test_think now run. * Load a checkpoint whose sampler mlx-vlm refuses with think off rather than failing (D-058) * Renumber the generation decision D-059: PR #136 claims D-058 * Keep forked caches apart, cancel before the prefill and after the last block, serialize the stub generation tests, fix a DocC link (#137 review, CI) LayerCache.extended() gives the copy array objects of its own: assigning into a slice of an MLXArray changes that object, so two continuations of one prefill wrote into each other's buffer (the new forkedContinuations test fails without this). generate returns cancelled for a task cancelled before the prefill or during the last block. The stub generation suite makes MLX random state, so it now runs inside the serialized MLX suite; in parallel it took down CI's test process. DiffusionSampler's link to MLXRandom.RandomState is code text, which DocC can resolve. * Serve chat on the mlx backend: DiffusionGemmaRuntime conforms to TextGenerator (#53, D-059) The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10. * Stream whole blocks: size the chat queue from the generator's block; skip only the channel markers in chat; regenerate the generation oracle with a long prompt and escaped dashes; tighten the generation tests (#137 review) * Docs, D-059 items 10 to 12, the long-stream live test and the measured figures (#137 review) * Merge main; drop two ambiguous DocC links to the stream's capacity
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Refs #53.
The HTTP and core side of
POST /v1/chat/completions: upstream'schat.pyfor the in-process generator (MlxGenerator), built and tested against a stub generator while the model side (#50, #51, #52) is ported. Nothing underSources/OpenJevDiffusionGemma/ModelorRuntimechanges. The issue closes when the model is wired.What was built
Generation/: theTextGeneratorprotocol andTextGeneration(below);ChatCompletionRequest, the route's checks andGenerator.normalize(passthrough fields only,max_tokens/max_completion_tokensa positive integer with a bool refused, default 1024, capped atOPENJEV_GEN_MAX_TOKENS, thinking off unless set, streamed usage forced on, JSON mode's instruction with the schema text);ChatCompletions, upstream's capacity (OPENJEV_GEN_MAX_INFLIGHTplusOPENJEV_GEN_MAX_QUEUE, then the 529overloaded_errorwithretry-after: 2; a request counts from the check, before its prompt renders, which also bounds the prompts rendering at once; a cancelled wait gives its place back), the prompt limit's 400 before the answer starts, single-token stop strings as stop ids, the thought-channel markers asskipSpecialTokenIDs, and whole replies withusageandextract_json;ChatCompletionStream, the queue of 64 pieces betweenemitand the response;ExtractJSON, CPython'sraw_decodesemantics; the reply and SSE event shapes;ChatCompletionError, OpenAI's{"error": {"message", "type", "code"}}.SystemOneServicegainstextGenerator(nilby default);DecisionEnginereturns its backend when the backend conforms.ChatCompletionsRoute, registered only when the service's model is aTextGenerator. The body is read asrequest.json()reads it. Streams aretext/event-stream; charset=utf-8: the role chunk, the content chunks, the finish chunk, the usage chunk,data: [DONE]. A sibling task watches the connection (ClientDisconnectHandler,ConnectionRegistry): a client that disconnects, or one so slow that 64 pieces wait, stops the generation at the next block, and a live client never loses a piece. A whole reply's client that leaves cancels it too (499 in the log). Generator failures are a 503 before the answer starts and are logged.SwiftTransformersTokenizer.generationPromptIDs(messages:thinking:), upstream'sprompt_ids(the chat template over the messages withadd_generation_promptandenable_thinking, then the empty thought scaffold<|channel>thought\n<channel|>), rendered on a thread of its own with 16 MB of stack; the template environment gains jinja2'ssequencetest, Python'sstr()intrim(Chat prompts trim a state as Foundation does, not as Python's str.strip: U+001C to U+001F stay, U+200B goes #124'strimstays) and jinja2'sdictsortorder (str.lower(), then code point, final sigma included), where swift-jinja's own compares keys with Foundation's case-insensitivecompare.Tools/fixtures/chat_tables.py, run bymake fixtures, records upstream's ownMlxGenerator.prompt_ids(51 conversations: system, user and assistant turns, multi-turn, text parts, thinking on and off, non-ASCII,trimcases, tool calls, argument keys the twodictsorts order differently, JSON mode),Generator.normalize(56 bodies),extract_json(44 replies) and 69 HTTP exchanges through upstream's route over the stub runtimes of itstest_mlx_backend.pyintoFixtures/chat-completions/(tokenizer files only, from the Hugging Face cache; no weights).test_a_slow_reader_cancels_rather_than_loses_chunks,test_a_one_token_reply_still_streams,test_a_cancelled_wait_for_a_chat_slot_leaks_no_capacityand the normalization andextract_jsonparity; in OpenJevServerTests the 69 recorded exchanges byte for byte,test_chat_normalizes_jev_ultrafast_request,test_chat_max_tokens_must_be_an_integer,test_chat_model_is_required,test_chat_rejects_an_unknown_model,test_chat_refuses_an_over_long_prompt,test_chat_capacity_is_refused,test_the_thought_channel_never_reaches_a_chat_client,test_the_thought_channel_never_reaches_a_streaming_client,test_a_one_token_reply_still_streams,test_chat_json_mode_and_max_tokens,test_chat_stream_on_mlx(stub),test_generation_model_is_listed, and on a live servertest_a_disconnected_client_stops_generation, a whole reply's client leaving, a client leaving while it waits, and graceful and timed-out shutdowns with a stream in flight; in OpenJevDiffusionGemmaTests the prompt parity (opt-in like the other tokenizer tests). Still disabled, naming Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51 and the wiring after this PR:test_chat_completion_on_mlx,test_chat_completion_generates_text,test_chat_stream_matches_the_whole_reply,test_chat_json_mode_returns_one_object,test_no_reply_leaks_the_thought_channel, and the live suite'stest_chatandtest_chat_stream; the DiffusionGemma ones have bodies that run throughSystemOneService.textGeneratoronce the runtime conforms.OPENJEV_GEN_*rows, docs/compatibility.md (chat: the route yes, the model Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51), docs/09, docs/development.md, Fixtures and Tools READMEs, THIRD_PARTY.md, README, and D-058.The protocol
generateis exactly upstream'sMlxRuntime.generatecontract:emitper committed token with the detokenizer's text, once more at the end with the final segment and aniltoken, andfalse(or cancelling the task) stops at the next block boundary. For the model side:DiffusionGemmaRuntimeconforms with the prompt, encoding, markers and limitnonisolated(generationPromptIDscan forward to the tokenizer's,thoughtChannelMarkerIDsisthoughtOpen + thoughtClose,[100, 45518, 107, 101]), and themlxbackend's chat route then appears throughDecisionEngine.textGeneratorwith nothing else to wire.generationPromptIDsisasyncbecause rendering a request's tool calls needs more stack than a task's thread has (below).Departures (D-058)
Generator's proxy to vLLM, and its two tests, are out of scope: this port has no vLLM backend.role; achat_template_kwargs,response_format,response_format.json_schemaor streamingstream_optionsthat is true but not a dict; astopof another type (upstream crashes after the prompt, a stream after its first chunk); a template error; messages nested past 64 levels. In JSON mode upstream answers a string message withdict()'s message; the port gives its own.api_errornaming the error's type withretry-after: 2, logged, where upstream's MLX generator gives a bare 500; after a stream has started, the connection closes without the response's end, as upstream's does, and it is logged.nullthe template writes with{{ }}(a tool call argument, a tool's missing result) is written as nothing where jinja2 writesNone, and an integer pastIntis a float; both only in tool calls, which upstream's MLX path never declares. An object with two keys that differ only in Unicode normalization (écomposed and decomposed) is a 400: Python keeps both keys, but swift-jinja keys objects by Swift'sString, so one would replace the other unseen. The 51 recorded conversations match text and ids exactly.extract_json: a lone surrogate escape is U+FFFD and a value nested past 1,024 levels leaves the reply unchanged, where upstream answers a 500.json.loadsreads (NaN,Infinity, lone surrogates): the 400 "The request body is not valid JSON." (D-016), where upstream serves them.Of the 69 recorded exchanges, 61 match upstream byte for byte, event streams included; the other 8 are these departures, which the replay test asserts explicitly.
Verification
All on Apple silicon with Swift 6.4 (Xcode), the branch rebased on
fad9e1a:make lint,swift build --build-testsandmake docs(DocC, warnings as errors) pass.make fixtureswith the pinned venv regenerates every file underFixtures/byte for byte;chat_tables.pygave identical files on 12 consecutive runs.swift test, DiffusionGemma model tests included (OPENJEV_MLX_CACHE_LIMIT_GB=4, checkpoint from the Hugging Face cache): 8 runs, 797 tests, all passed in about 11 minutes;.github/scripts/check-test-log.shpasses on its log (onlyOPENJEV_*opt-in skips).swift build -c release -Xswiftc -enable-testing, run throughswiftpm-testing-helper): OpenJevCoreTests 20 and OpenJevServerTests 23, three runs each, all passed, for Swift 6.4's-Otask-group miscompile (Serve JevK5 (jevk5-0.2) on MLX, with its conversions and parity against the author's run (D-052) #122).make lint,swift build --build-tests,make docsandmake fixtures(only the two new prompt rows change) pass; OpenJevCoreTests (260), OpenJevServerTests (139) and the tokenizer suites (30) pass withcheck-test-log.sh; the chat tests in a release build (OpenJevCoreTests 24, OpenJevServerTests 24) passed three runs each, the release build itself without Swift 6.4's CoroSplit crash. Each new test was checked to fail with its fix reverted. The full suite was not rerun for this round, which changes nothing the model tests exercise.Found
run_in_executorregisters its callback completes the awaited future at once, ahead of the pieces queued withcall_soon_threadsafe. Its own one-token stream lost its only piece on a fresh app in 16 of 30 tries; 3.12 (its container) is unaffected and a real model never finishes that fast. The fixture script works around it for determinism; the port has no such ordering.[100, 45518, 107, 101]holds two ordinary tokens,thoughtand\n. If mlx-vlm'sskip_special_token_idsdrops them everywhere, chat replies lose those tokens; Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51 should check that against mlx-vlm (not installed here, so not confirmed).dictsortorders keys with Foundation's case-insensitivecompare, not jinja2'sstr.lower()and code point:ßsorts asssand a decomposedéwith a composed one. Upstream's own rendering of tool call arguments with such keys is now in the fixtures, and the port matches it.