Skip to content

Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) - #136

Merged
alaineid merged 3 commits into
mainfrom
issue-53-chat-completions
Oct 6, 2026
Merged

alaineid merged 3 commits into
mainfrom
issue-53-chat-completions

Conversation

@alaineid

@alaineid alaineid commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Refs #53.

The HTTP and core side of POST /v1/chat/completions: upstream's chat.py for the in-process generator (MlxGenerator), built and tested against a stub generator while the model side (#50, #51, #52) is ported. Nothing under Sources/OpenJevDiffusionGemma/Model or Runtime changes. The issue closes when the model is wired.

What was built

  • OpenJevCore, Generation/: the TextGenerator protocol and TextGeneration (below); ChatCompletionRequest, the route's checks and Generator.normalize (passthrough fields only, max_tokens/max_completion_tokens a positive integer with a bool refused, default 1024, capped at OPENJEV_GEN_MAX_TOKENS, thinking off unless set, streamed usage forced on, JSON mode's instruction with the schema text); ChatCompletions, upstream's capacity (OPENJEV_GEN_MAX_INFLIGHT plus OPENJEV_GEN_MAX_QUEUE, then the 529 overloaded_error with retry-after: 2; a request counts from the check, before its prompt renders, which also bounds the prompts rendering at once; a cancelled wait gives its place back), the prompt limit's 400 before the answer starts, single-token stop strings as stop ids, the thought-channel markers as skipSpecialTokenIDs, and whole replies with usage and extract_json; ChatCompletionStream, the queue of 64 pieces between emit and the response; ExtractJSON, CPython's raw_decode semantics; the reply and SSE event shapes; ChatCompletionError, OpenAI's {"error": {"message", "type", "code"}}. SystemOneService gains textGenerator (nil by default); DecisionEngine returns its backend when the backend conforms.
  • OpenJevServer: ChatCompletionsRoute, registered only when the service's model is a TextGenerator. The body is read as request.json() reads it. Streams are text/event-stream; charset=utf-8: the role chunk, the content chunks, the finish chunk, the usage chunk, data: [DONE]. A sibling task watches the connection (ClientDisconnectHandler, ConnectionRegistry): a client that disconnects, or one so slow that 64 pieces wait, stops the generation at the next block, and a live client never loses a piece. A whole reply's client that leaves cancels it too (499 in the log). Generator failures are a 503 before the answer starts and are logged.
  • OpenJevDiffusionGemma tokenizer: SwiftTransformersTokenizer.generationPromptIDs(messages:thinking:), upstream's prompt_ids (the chat template over the messages with add_generation_prompt and enable_thinking, then the empty thought scaffold <|channel>thought\n<channel|>), rendered on a thread of its own with 16 MB of stack; the template environment gains jinja2's sequence test, Python's str() in trim (Chat prompts trim a state as Foundation does, not as Python's str.strip: U+001C to U+001F stay, U+200B goes #124's trim stays) and jinja2's dictsort order (str.lower(), then code point, final sigma included), where swift-jinja's own compares keys with Foundation's case-insensitive compare.
  • Fixtures: Tools/fixtures/chat_tables.py, run by make fixtures, records upstream's own MlxGenerator.prompt_ids (51 conversations: system, user and assistant turns, multi-turn, text parts, thinking on and off, non-ASCII, trim cases, tool calls, argument keys the two dictsorts order differently, JSON mode), Generator.normalize (56 bodies), extract_json (44 replies) and 69 HTTP exchanges through upstream's route over the stub runtimes of its test_mlx_backend.py into Fixtures/chat-completions/ (tokenizer files only, from the Hugging Face cache; no weights).
  • Tests, named after upstream's: in OpenJevCoreTests test_a_slow_reader_cancels_rather_than_loses_chunks, test_a_one_token_reply_still_streams, test_a_cancelled_wait_for_a_chat_slot_leaks_no_capacity and the normalization and extract_json parity; in OpenJevServerTests the 69 recorded exchanges byte for byte, test_chat_normalizes_jev_ultrafast_request, test_chat_max_tokens_must_be_an_integer, test_chat_model_is_required, test_chat_rejects_an_unknown_model, test_chat_refuses_an_over_long_prompt, test_chat_capacity_is_refused, test_the_thought_channel_never_reaches_a_chat_client, test_the_thought_channel_never_reaches_a_streaming_client, test_a_one_token_reply_still_streams, test_chat_json_mode_and_max_tokens, test_chat_stream_on_mlx (stub), test_generation_model_is_listed, and on a live server test_a_disconnected_client_stops_generation, a whole reply's client leaving, a client leaving while it waits, and graceful and timed-out shutdowns with a stream in flight; in OpenJevDiffusionGemmaTests the prompt parity (opt-in like the other tokenizer tests). Still disabled, naming Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51 and the wiring after this PR: test_chat_completion_on_mlx, test_chat_completion_generates_text, test_chat_stream_matches_the_whole_reply, test_chat_json_mode_returns_one_object, test_no_reply_leaks_the_thought_channel, and the live suite's test_chat and test_chat_stream; the DiffusionGemma ones have bodies that run through SystemOneService.textGenerator once the runtime conforms.
  • Docs: docs/05 (the chat route), docs/deployment.md (the generation settings, the 529, a text-generation section), the configuration reference's OPENJEV_GEN_* rows, docs/compatibility.md (chat: the route yes, the model Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51), docs/09, docs/development.md, Fixtures and Tools READMEs, THIRD_PARTY.md, README, and D-058.

The protocol

public protocol TextGenerator: Sendable {
    var maxPromptTokens: Int { get }
    var thoughtChannelMarkerIDs: [Int] { get }
    func generationPromptIDs(messages: [JSONValue], thinking: Bool) async throws -> [Int]
    func encode(_ text: String) throws -> [Int]
    func generate(
        prompt: [Int], maxTokens: Int, stopIDs: [Int], skipSpecialTokenIDs: [Int],
        emit: @Sendable (_ text: String, _ token: Int?) -> Bool
    ) async throws -> TextGeneration
}

public struct TextGeneration: Sendable, Hashable {
    public enum FinishReason: String, Sendable, Hashable { case stop, length, cancelled }
    public var generated: [Int]      // without the stop token
    public var promptTokens: Int
    public var finishReason: FinishReason
}

generate is exactly upstream's MlxRuntime.generate contract: emit per committed token with the detokenizer's text, once more at the end with the final segment and a nil token, and false (or cancelling the task) stops at the next block boundary. For the model side: DiffusionGemmaRuntime conforms with the prompt, encoding, markers and limit nonisolated (generationPromptIDs can forward to the tokenizer's, thoughtChannelMarkerIDs is thoughtOpen + thoughtClose, [100, 45518, 107, 101]), and the mlx backend's chat route then appears through DecisionEngine.textGenerator with nothing else to wire. generationPromptIDs is async because rendering a request's tool calls needs more stack than a task's thread has (below).

Departures (D-058)

  • No vLLM proxy. Generator's proxy to vLLM, and its two tests, are out of scope: this port has no vLLM backend.
  • A 400 where upstream crashes with a bare 500: a message that is not an object or has no string role; a chat_template_kwargs, response_format, response_format.json_schema or streaming stream_options that is true but not a dict; a stop of another type (upstream crashes after the prompt, a stream after its first chunk); a template error; messages nested past 64 levels. In JSON mode upstream answers a string message with dict()'s message; the port gives its own.
  • Generator failures: a 503 api_error naming the error's type with retry-after: 2, logged, where upstream's MLX generator gives a bare 500; after a stream has started, the connection closes without the response's end, as upstream's does, and it is logged.
  • Streaming: a piece that finds the queue full ends the reply after every queued piece (upstream's end marker displaced the oldest one, a gap); a disconnect is seen from the connection at once (upstream polls every 0.1 s); a whole reply's client that leaves cancels it (upstream runs it to the end).
  • Capacity: a request counts from the capacity check, before its prompt renders. Upstream counts it once the prompt has rendered, but renders on its event loop, so nothing runs in between; here prompts render concurrently on threads of their own, and counting first bounds them by the configured limit. A request refused afterwards (a 400 for its prompt, say) holds a place while it renders.
  • Prompt rendering: a null the template writes with {{ }} (a tool call argument, a tool's missing result) is written as nothing where jinja2 writes None, and an integer past Int is a float; both only in tool calls, which upstream's MLX path never declares. An object with two keys that differ only in Unicode normalization (é composed and decomposed) is a 400: Python keeps both keys, but swift-jinja keys objects by Swift's String, so one would replace the other unseen. The 51 recorded conversations match text and ids exactly.
  • extract_json: a lone surrogate escape is U+FFFD and a value nested past 1,024 levels leaves the reply unchanged, where upstream answers a 500.
  • Bodies only json.loads reads (NaN, Infinity, lone surrogates): the 400 "The request body is not valid JSON." (D-016), where upstream serves them.

Of the 69 recorded exchanges, 61 match upstream byte for byte, event streams included; the other 8 are these departures, which the replay test asserts explicitly.

Verification

All on Apple silicon with Swift 6.4 (Xcode), the branch rebased on fad9e1a:

  • make lint, swift build --build-tests and make docs (DocC, warnings as errors) pass.
  • make fixtures with the pinned venv regenerates every file under Fixtures/ byte for byte; chat_tables.py gave identical files on 12 consecutive runs.
  • The full swift test, DiffusionGemma model tests included (OPENJEV_MLX_CACHE_LIMIT_GB=4, checkpoint from the Hugging Face cache): 8 runs, 797 tests, all passed in about 11 minutes; .github/scripts/check-test-log.sh passes on its log (only OPENJEV_* opt-in skips).
  • The chat tests in a release build (swift build -c release -Xswiftc -enable-testing, run through swiftpm-testing-helper): OpenJevCoreTests 20 and OpenJevServerTests 23, three runs each, all passed, for Swift 6.4's -O task-group miscompile (Serve JevK5 (jevk5-0.2) on MLX, with its conversions and parity against the author's run (D-052) #122).
  • After the first review round (Copilot's three comments and an independent review): make lint, swift build --build-tests, make docs and make fixtures (only the two new prompt rows change) pass; OpenJevCoreTests (260), OpenJevServerTests (139) and the tokenizer suites (30) pass with check-test-log.sh; the chat tests in a release build (OpenJevCoreTests 24, OpenJevServerTests 24) passed three runs each, the release build itself without Swift 6.4's CoroSplit crash. Each new test was checked to fail with its fix reverted. The full suite was not rerun for this round, which changes nothing the model tests exercise.
  • The debug chat suites passed eight runs in a row; the new shutdown test fails with the cut-short fix removed.
  • Not run locally: the Linux job (Swift 6.2). CI covers it.

Found

  • Upstream's stream can lose its last pieces under CPython 3.14: a call that finishes before run_in_executor registers its callback completes the awaited future at once, ahead of the pieces queued with call_soon_threadsafe. Its own one-token stream lost its only piece on a fresh app in 16 of 30 tries; 3.12 (its container) is unaffected and a real model never finishes that fast. The fixture script works around it for determinism; the port has no such ordering.
  • Upstream's chat skip list [100, 45518, 107, 101] holds two ordinary tokens, thought and \n. If mlx-vlm's skip_special_token_ids drops them everywhere, chat replies lose those tokens; Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51 should check that against mlx-vlm (not installed here, so not confirmed).
  • In a debug build swift-jinja overflows a task's 512 KB stack on a tool call nested about 16 levels, which any client could send: hence the 16 MB rendering thread and the 64-level bound.
  • swift-jinja 2.5.1's dictsort orders keys with Foundation's case-insensitive compare, not jinja2's str.lower() and code point: ß sorts as ss and a decomposed é with a composed one. Upstream's own rendering of tool call arguments with such keys is now in the fixtures, and the port matches it.

…de of #53 (D-058)

Port upstream's chat.py for the in-process generator (MlxGenerator), against a stub, while the
model's generation is ported in parallel (#50 to #52). OpenJevCore declares TextGenerator, whose
generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) is MlxRuntime.generate, with the
prompt, stop encoding, thought-channel markers and prompt limit the route reads; the server adds
POST /v1/chat/completions when SystemOneService.textGenerator is set, which DecisionEngine gives
for a backend that conforms. The route checks, normalizes, bounds capacity (529, retry-after 2),
refuses an over-long prompt before the answer starts, answers whole replies with usage and JSON
mode's extraction, and streams OpenAI's chunks through a queue of 64 pieces; a client that goes
away, or a reader that falls 64 pieces behind, stops the generation at its next block, and a reply
that ends early is always a prefix.

SwiftTransformersTokenizer renders a chat request's messages as upstream's prompt_ids does, the
scaffold after, on a 16 MB thread, with jinja2's sequence test and Python's str() in trim.
Tools/fixtures/chat_tables.py records upstream's normalize, extract_json, prompt_ids and 69 route
exchanges (Fixtures/chat-completions); the port matches 61 byte for byte, event streams included,
and the other 8 are the departures D-058 records. The tests carry upstream's names; the ones that
need the model stay disabled, naming #51.

Refs #53.
Item 6 adds how a graceful and a timed-out shutdown treat a stream in flight, which the
route now handles and tests; item 11 says which disabled cases have bodies and reflows two
paragraphs.

Refs #53.
Copilot AI balanced review requested due to automatic review settings October 6, 2026 14:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unbounded rendering threads, Unicode-equivalent key loss, and an unlogged backend failure path need correction.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity · 1 Low severity

Open (3)
What changed in this PR

Adds the core and HTTP infrastructure for OpenAI-compatible chat completions, ready for DiffusionGemma generation wiring.

Changes:

  • Introduces request normalization, generation capacity, streaming, JSON extraction, and response/error contracts.
  • Adds conditional server routing, cancellation handling, tokenizer prompt rendering, and fixture-backed parity tests.
  • Documents chat configuration, compatibility, architecture, and deployment behavior.
File Description
Sources/​OpenJevCore/​Generation/​TextGenerator.swift Defines generation protocol and result contract.
Sources/​OpenJevCore/​Generation/​ChatCompletionRequest.swift Validates and normalizes chat requests.
Sources/​OpenJevCore/​Generation/​ChatCompletions.swift Implements preparation, capacity, and completion.
Sources/​OpenJevCore/​Generation/​ChatCompletionStream.swift Implements bounded SSE streaming.
Sources/​OpenJevCore/​Generation/​ChatCompletionResponse.swift Defines replies, usage, and SSE events.
Sources/​OpenJevCore/​Generation/​ChatCompletionError.swift Defines OpenAI-shaped errors.
Sources/​OpenJevCore/​Generation/​ExtractJSON.swift Extracts JSON-mode responses.
Sources/​OpenJevCore/​Engine/​SystemOneService.swift Exposes optional text generation.
Sources/​OpenJevCore/​JSON/​JSONValue+Python.swift Adds Python truthiness and representations.
Sources/​OpenJevCore/​JSON/​PythonJSONWriter.swift Adds non-finite-number output support.
Sources/​OpenJevCore/​Schema/​TextOf.swift Adds Python-compatible right trimming.
Sources/​OpenJevServer/​ChatCompletionsRoute.swift Serves whole and streamed completions.
Sources/​OpenJevServer/​OpenJevApplication.swift Conditionally registers the chat route.
Sources/​OpenJevServer/​WireResponses.swift Renders chat errors.
Sources/​OpenJevServer/​RefusalLog.swift Logs generation failures.
Sources/​OpenJevServer/​Documentation.docc/​Configuration.md Documents generation settings.
Sources/​OpenJevDiffusionGemma/​Tokenization/​SwiftTransformersTokenizer.swift Builds generation prompts and Jinja parity behavior.
Sources/​OpenJevTestSupport/​StubTextGenerator.swift Supplies stub generation runtimes.
Sources/​OpenJevTestSupport/​StubGeneratingBackend.swift Combines stub reads and generation.
Sources/​OpenJevTestSupport/​ChatFixtures.swift Loads recorded chat fixtures.
Tests/​OpenJevCoreTests/​Generation/​ChatCompletionsTests.swift Tests capacity, streaming, and replies.
Tests/​OpenJevCoreTests/​Generation/​ChatFixtureTests.swift Tests normalization and JSON extraction parity.
Tests/​OpenJevServerTests/​Support/​ChatHarness.swift Provides chat route test infrastructure.
Tests/​OpenJevServerTests/​ChatCompletionsRouteTests.swift Tests the HTTP contract.
Tests/​OpenJevServerTests/​ChatRouteFixtureTests.swift Replays recorded exchanges.
Tests/​OpenJevServerTests/​ChatLiveServerTests.swift Tests live cancellation and shutdown.
Tests/​OpenJevDiffusionGemmaTests/​Tokenization/​ChatTemplateTrimTests.swift Tests Python/Jinja semantics.
Tests/​OpenJevDiffusionGemmaTests/​Tokenization/​ChatCompletionPromptTests.swift Tests prompt parity and nesting.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​UpstreamReadCaseTests.swift Adds disabled model-backed chat cases.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​RuntimeTests.swift Verifies generation remains unwired.
Tests/​OpenJevLiveTests/​LiveTests.swift Updates live-test prerequisites.
Tools/​fixtures/​chat_tables.py Generates upstream chat fixtures.
Tools/​README.md Documents the fixture generator.
Makefile Runs chat fixture generation.
Fixtures/​README.md Catalogues chat fixtures.
Fixtures/​chat-completions/​README.md Documents fixture contents.
Fixtures/​chat-completions/​prompts.json Records prompt parity cases.
Fixtures/​chat-completions/​normalize.json Records normalization cases.
Fixtures/​chat-completions/​extract_json.json Records JSON extraction cases.
Fixtures/​chat-completions/​routes.json Records HTTP exchanges.
README.md Updates chat availability status.
THIRD_PARTY.md Records CPython-derived scanner behavior.
docs/​05-architecture.md Describes chat architecture.
docs/​06-decisions.md Records D-058 decisions.
docs/​09-conformance-and-testing.md Documents parity coverage.
docs/​compatibility.md Updates compatibility claims.
docs/​deployment.md Documents chat operation and settings.
docs/​development.md Updates fixture commands.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread Sources/OpenJevDiffusionGemma/Tokenization/SwiftTransformersTokenizer.swift Outdated
Comment thread Sources/OpenJevServer/ChatCompletionsRoute.swift Outdated
…ja2 does

The first review of PR #136 (Copilot and an independent review):

- ChatCompletions.prepare counts a request against the capacity bound at the check, before its
  prompt renders, and gives the place back when it is refused afterwards. Upstream counts after
  rendering, but renders on its event loop; here each prompt renders on a 16 MB thread of its
  own, which a burst could otherwise start without limit. PreparedChatCompletion carries the
  place and is answered once. The bound saturates at Int.max instead of trapping.
- The route logs the 503 prepare gives for a stop string the generator cannot encode, as it logs
  a generation's.
- The chat template's dictsort is jinja2's: keys ordered by Python's str.lower(), final sigma
  included, then code point, where swift-jinja's compares them with Foundation's
  case-insensitive compare. Two prompts.json rows record upstream's rendering of such keys.
- An object whose keys differ only in Unicode normalization is a 400: swift-jinja keys objects
  by Swift's String and would keep one of them.
- The live disconnect test waits for the next request's 200 rather than the stub's count, which
  drops before the stream gives its place back.

Refs #53.
@alaineid
alaineid merged commit d25223a into main Oct 6, 2026
9 checks passed
@alaineid
alaineid deleted the issue-53-chat-completions branch October 6, 2026 15:18
alaineid added a commit that referenced this pull request Oct 6, 2026
…Generator (#53, D-059)

The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10.
alaineid added a commit that referenced this pull request Oct 6, 2026
…l then" (#140)

PR #136 added `extension ChatCompletionsConfiguration` with an `init(_:)` in
Sources/OpenJevServer/ChatCompletionsRoute.swift, next to the extensions of EngineConfiguration
and EncoderEngineConfiguration in BackendProvider.swift. Those are the only three extensions in
Sources/OpenJevServer, and OpenJevCore declares all three types. docs/development.md's Links
bullet and the header of Tools/docs/build-site.sh still said two; both say three now, and
development.md names the three `init(_:)` the server adds.

GitHub Pages publishes from GitHub Actions: the Pages API reports build_type workflow, and the
Documentation run on main for d25223a deployed the site. The header of
.github/workflows/docs.yml now says the deploy job is skipped "otherwise" instead of "until
then", as PR #139 words docs/development.md, whose own copy of that line PR #139 changes.
docs/06-decisions.md keeps its wording: a decision records what was true when it was written.
alaineid added a commit that referenced this pull request Oct 6, 2026
Every public symbol declared in the five library modules now has a doc comment, as the Coverage
bullet of docs/development.md says. Comments only: no code, signature or access level changes.

- VisionTower.swift (#129): the @ModuleInfo and @ParameterInfo properties and the callAsFunction
  methods of VisionAttention, VisionMLP, VisionBlock, VisionPatchEmbedder, VisionModel,
  VisionEncoder and MultimodalEmbedder (29), and the callAsFunction of ClippableLinear and
  VisionRMSNorm. A property says what it holds and its checkpoint key, as the text model's do; a
  callAsFunction says what it computes, citing mlx-vlm's __call__, with the file's shapes.
- VisionError.description (#121): "The message.", as the other errors' descriptions say.
- extension ChatCompletionsConfiguration in OpenJevServer (#136): a /// line on the extension,
  as BackendProvider.swift's two extensions of core types have.
alaineid added a commit that referenced this pull request Oct 6, 2026
…er, the block loop, think and the chat model (#50, #51, #52, #53, D-059) (#137)

* Generate text and think on the MLX backend: the diffusion sampler, the block loop and the think option (#50, #51, #52, D-058)

Port mlx-vlm 0.6.15's sampling functions (DiffusionSampler), one block of
stream_diffusion_generate as upstream's MlxRuntime.generate runs it
(greedy, the default confidence-threshold sampler, the linear temperature
schedule, self-conditioning, stable-and-confident stopping), and
diffusion_update_cache on copies of the layer caches. DiffusionGemmaRuntime
gains generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) with
canvas sizing, EOS and stop ids, the SentencePiece streaming detokenizer,
emit and task cancellation, and think(prompt:budget:stopIDs:) on it, so
capabilities.think is on. Each reply's canvases come from a RandomState
seeded with Configuration.generationSeed (0), which upstream leaves
unseeded.

Tools/fixtures/generation_oracle.py records the oracle in
Fixtures/generation: sampler vectors, seven seeded greedy replies block by
block, and two think requests. On the oracle's Metal library every reply
and thought matches token for token with upstream's billing; natively the
short replies match whole and the long story for its first 8 tokens.
Upstream's three think cases and the live suite's test_think now run.

* Load a checkpoint whose sampler mlx-vlm refuses with think off rather than failing (D-058)

* Renumber the generation decision D-059: PR #136 claims D-058

* Keep forked caches apart, cancel before the prefill and after the last block, serialize the stub generation tests, fix a DocC link (#137 review, CI)

LayerCache.extended() gives the copy array objects of its own: assigning into a slice of an MLXArray changes that object, so two continuations of one prefill wrote into each other's buffer (the new forkedContinuations test fails without this). generate returns cancelled for a task cancelled before the prefill or during the last block. The stub generation suite makes MLX random state, so it now runs inside the serialized MLX suite; in parallel it took down CI's test process. DiffusionSampler's link to MLXRandom.RandomState is code text, which DocC can resolve.

* Serve chat on the mlx backend: DiffusionGemmaRuntime conforms to TextGenerator (#53, D-059)

The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10.

* Stream whole blocks: size the chat queue from the generator's block; skip only the channel markers in chat; regenerate the generation oracle with a long prompt and escaped dashes; tighten the generation tests (#137 review)

* Docs, D-059 items 10 to 12, the long-stream live test and the measured figures (#137 review)

* Merge main; drop two ambiguous DocC links to the stream's capacity
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants