Repository navigation
Generate text, think and chat on the MLX backend: the diffusion sampler, the block loop, think and the chat model (#50, #51, #52, #53, D-059) - #137
Merged
Conversation
…e block loop and the think option (#50, #51, #52, D-058) Port mlx-vlm 0.6.15's sampling functions (DiffusionSampler), one block of stream_diffusion_generate as upstream's MlxRuntime.generate runs it (greedy, the default confidence-threshold sampler, the linear temperature schedule, self-conditioning, stable-and-confident stopping), and diffusion_update_cache on copies of the layer caches. DiffusionGemmaRuntime gains generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) with canvas sizing, EOS and stop ids, the SentencePiece streaming detokenizer, emit and task cancellation, and think(prompt:budget:stopIDs:) on it, so capabilities.think is on. Each reply's canvases come from a RandomState seeded with Configuration.generationSeed (0), which upstream leaves unseeded. Tools/fixtures/generation_oracle.py records the oracle in Fixtures/generation: sampler vectors, seven seeded greedy replies block by block, and two think requests. On the oracle's Metal library every reply and thought matches token for token with upstream's billing; natively the short replies match whole and the long story for its first 8 tokens. Upstream's three think cases and the live suite's test_think now run.
… than failing (D-058)
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Cache-copy isolation and cancellation handling have unresolved correctness issues.
Review effort: Balanced
Findings: 1
Open (4)
What changed in this PR
Adds MLX text generation and enables think, providing the model-side foundation for the future chat endpoint.
Changes:
- Implements sampling, block denoising, cache continuation and streaming text.
- Enables seeded thought generation through the runtime.
- Adds upstream oracle comparisons, tests and documentation.
| File | Description |
|---|---|
| Tools/README.md | Lists the generation oracle tool. |
| Tools/oracle/results/generation_run.json | Records oracle timings and reproducibility. |
| Tools/fixtures/generation_oracle.py | Generates sampling, reply and thought fixtures. |
| Tests/OpenJevLiveTests/LiveTests.swift | Enables live think testing. |
| Tests/OpenJevDiffusionGemmaTests/Support/GenerationFixtures.swift | Decodes generation fixtures. |
| Tests/OpenJevDiffusionGemmaTests/Runtime/UpstreamReadCaseTests.swift | Enables upstream think cases. |
| Tests/OpenJevDiffusionGemmaTests/Runtime/RuntimeTests.swift | Updates capability and refusal tests. |
| Tests/OpenJevDiffusionGemmaTests/Runtime/RuntimeStubs.swift | Adds scripted generation support. |
| Tests/OpenJevDiffusionGemmaTests/Runtime/ImageRuntimeTests.swift | Checks image/think refusals. |
| Tests/OpenJevDiffusionGemmaTests/Generation/StreamingDetokenizerTests.swift | Tests streamed text and byte handling. |
| Tests/OpenJevDiffusionGemmaTests/Generation/SamplerTests.swift | Checks sampling against oracle vectors. |
| Tests/OpenJevDiffusionGemmaTests/Generation/GenerationRuntimeTests.swift | Tests block-loop behavior without weights. |
| Tests/OpenJevDiffusionGemmaTests/Generation/GenerationOracleTests.swift | Compares model replies and thoughts upstream. |
| Sources/OpenJevDiffusionGemma/Runtime/RuntimeConfiguration.swift | Adds the generation seed. |
| Sources/OpenJevDiffusionGemma/Runtime/Generation.swift | Implements generation and think. |
| Sources/OpenJevDiffusionGemma/Runtime/DiffusionGemmaRuntime.swift | Wires generation and capabilities. |
| Sources/OpenJevDiffusionGemma/Model/Prefill.swift | Adds committed-block cache updates. |
| Sources/OpenJevDiffusionGemma/Model/ModelTree.swift | Supplies sliding-cache window sizes. |
| Sources/OpenJevDiffusionGemma/Model/LayerCache.swift | Supports cache copying and extension. |
| Sources/OpenJevDiffusionGemma/Model/Configuration.swift | Uses Double configuration values. |
| Sources/OpenJevDiffusionGemma/Generation/StreamingDetokenizer.swift | Converts tokens into streamed text. |
| Sources/OpenJevDiffusionGemma/Generation/DiffusionSampler.swift | Implements upstream sampling functions. |
| Sources/OpenJevDiffusionGemma/Generation/BlockDenoising.swift | Adds generation policy and block denoising. |
| Sources/OpenJevDiffusionGemma/Documentation.docc/ReadingWithDiffusionGemma.md | Documents generation usage. |
| Sources/OpenJevDiffusionGemma/Documentation.docc/OpenJevDiffusionGemma.md | Exposes generation documentation. |
| Fixtures/README.md | Lists generation fixtures. |
| Fixtures/generation/README.md | Explains oracle contents and seeding. |
| docs/compatibility.md | Updates think support and differences. |
| docs/09-conformance-and-testing.md | Documents generation validation tiers. |
| docs/07-risks-and-unknowns.md | Updates generation parity risks. |
| docs/06-decisions.md | Records D-059 and intentional departures. |
| docs/05-architecture.md | Describes generation integration. |
| docs/03-diffusiongemma.md | Explains the upstream generation policy. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+73
to
+75
| let trim = keys.dim(2) - window + 1 | ||
| let keptKeys = trim > 0 ? keys[.ellipsis, trim..., 0...] : keys | ||
| let keptValues = trim > 0 ? values[.ellipsis, trim..., 0...] : values |
Contributor
Author
There was a problem hiding this comment.
Added in a99890b: LayerCacheTests.slidingTrim uses distinct positions and a window of 4. It updates twice past the window and checks the positions kept, the values and the offsets each time, then checks the prefill is unchanged after updating its extended() copy. The suite needs no checkpoint and runs in CI.
…t block, serialize the stub generation tests, fix a DocC link (#137 review, CI) LayerCache.extended() gives the copy array objects of its own: assigning into a slice of an MLXArray changes that object, so two continuations of one prefill wrote into each other's buffer (the new forkedContinuations test fails without this). generate returns cancelled for a task cancelled before the prefill or during the last block. The stub generation suite makes MLX random state, so it now runs inside the serialized MLX suite; in parallel it took down CI's test process. DiffusionSampler's link to MLXRandom.RandomState is code text, which DocC can resolve.
…tion-and-think # Conflicts: # Tests/OpenJevLiveTests/LiveTests.swift # docs/06-decisions.md # docs/09-conformance-and-testing.md # docs/compatibility.md
…Generator (#53, D-059) The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10.
This was referenced Oct 6, 2026
…skip only the channel markers in chat; regenerate the generation oracle with a long prompt and escaped dashes; tighten the generation tests (#137 review)
This was referenced Oct 6, 2026
This was referenced Oct 6, 2026
alaineid
added a commit
that referenced
this pull request
Oct 6, 2026
…act, and catch the configuration reference up with JevK5 (#147) Configuration.md's OPENJEV_GEN_MAX_INFLIGHT row said the chat route would serve once the model generates text (issue #51); PR #137 (D-059) made the mlx backend serve it. docs/10's encoder section said OpenJevCore will mirror upstream's EncoderEngine; QuestionReadBackend and EncoderDecisionEngine have done so since PR #82 (#67). Its batching bullet said "per forward pass", where the engine bounds each backend call and a backend may split a call into passes. Three Configuration.md sentences still described the port before PR #122 added JevK5: "Two are this port's own" (OPENJEV_JEVK5_MODEL makes three), "as upstream's do" (upstream's JevK5 keeps OPENJEV_JEVK5_WORKERS reads in flight), and JevK5's upstream variables listed under "the backends this port does not have". Docs only; the table's variable and default cells are unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Closes #50.
Closes #51.
Closes #52.
Closes #53.
Milestone 5's model half (#49): text generation on the MLX backend, the
thinkoption, and the chat endpoint (#136) wired to the model. Decision D-059; D-058 is the chat route's, from #136.What was built
Oracle.
Tools/fixtures/generation_oracle.pywritesFixtures/generation/generation.jsonfrom upstream's own path (MlxRuntime.generate,MlxEngine.decide) on mlx-vlm 0.6.15. It holds:diffusion.pylines 285 to 505, computed on the CPU;emitcall: a short answer, a list, JSON, a 640-token story over blocks of 256/256/128, a reply ended by an extra stop id, one cut bymax_tokens40, a 320-token report after a 1,235-token prompt, and one under another seed. The report's first commit trims the sliding layers past their 1,023-position window;thinkrequests (the README example withthink128, and 24 sequential nouls withthink64), with every answer after the thought.Everything runs twice, the second time reversed, and nothing is written unless both runs agree bit for bit. Dashes the model writes are stored as JSON escapes.
--checkreproduced the committed file, and the committed run record holds that check.Seeding. Upstream never seeds MLX's generator, so its replies vary from process to process. The port gives each reply
MLXRandom.RandomState(seed: Configuration.generationSeed)(0 by default), which draws whatmx.random.seedgives. That is tested on the sampler and on every block's recorded initial canvas.Diffusion sampler: entropy-bound canvas update, temperature schedule, stable-and-confident stopping #50.
DiffusionSampler, held to mlx-vlm's vectors bit for bit.Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51.
denoiseBlock,updateCache(diffusion_update_cache, on copies that never share array objects), the SentencePiece streaming detokenizer, andDiffusionGemmaRuntime.generate. Canvases aremin(256, max(remaining, 64)); it stops at EOS {1, 106, 50}, stop ids andmax_tokens; it callsemitper token plus a tail; a cancelled task ends the reply before the prefill, before the next block or after the last one; and it reuses a read's cached prefill.The think option: thought generation before a read, prefix continuation, billing #52.
thinkruns ongenerate, andcapabilities.thinkis on./v1/chat/completions: OpenAI-compatible generation with streaming, JSON mode and cancellation #53.
DiffusionGemmaRuntimeconforms to Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) #136'sTextGenerator, so the server servesPOST /v1/chat/completionsonmlx.Review fixes in this round
generateemits a committed block of up to 256 tokens back to back, and the chat stream's queue was upstream's 64. So the 65th piece of every block of more than 64 tokens ended the reply without its finish chunk or[DONE].max(64, 2 × blockLength + 1)pieces, 513 here, from a newTextGenerator.blockLengththat the generator reports (256 for the runtime). I chose a reported size over a constant because the block length is the model's: a token-by-token generator reports 1 and keeps 64.wholeBlockStreamsemits 257 pieces from one synchronous call; it failed withreaderFellBehindbefore the change and passes now. The slow-reader test holds at the new size.curl -N: a 451-token story streams to its finish and[DONE], equal to the same request withoutstream. The new live testlongStreamIsWholechecks this.[100, 45518, 107, 101]also drops the single newline token 107 andthought45518 wherever they appear, so lists run together and{"thought": 1}becomes{"": 1}.generate's parity tests keep upstream's list (D-059 item 10).teamandtoneincluded, probabilities and confidence: all bit for bit.commitsCopycompares the prefill's tensor digests after two continuations, short prompt and long.mlx_vlm_oracle.pydoes.Measured (2026-10-06, the reference Mac,
OPENJEV_MLX_CACHE_LIMIT_GB=4)Exact tier (
OPENJEV_MLX_METALLIB= the wheel's metallib, oracle RoPE table):emitcall.ReadOracleTests(the 63 reads, bit for bit) andSamplerTestspass.Native tier:
max_tokensThe long replies part where a near-tied argmax flips under mlx-swift's kernels. #51's long-reply agreement is therefore held in the exact tier, as D-014 holds reads (D-059 item 12).
Natively the seeded initial canvases match wherever the blocks before them drew alike. Both thoughts bill as upstream does (128/469 and 64/1,896), though their ids part after 48 and 3 tokens.
#52's "available for logging at debug level": the thought's ids are available to the engine, in the read's prefix, never in the
Decision.OpenJevCorehas no logger, so nothing logs them.Whole package, natively with the model tests on, run in groups because Xcode crashes on a single full run:
OPENJEV_JEVK5_MODEL), the Hub download (OPENJEV_TEST_DOWNLOAD), the vision full tensors, and the live suite withoutOPENJEV_LIVE_URL.encodersWithoutCoreML, which compiles only without Core ML..github/scripts/check-test-log.shpasses on all eight groups' logs.make lintis clean.OpenJevCore-iOSandOpenJevDiffusionGemmabuild for the iOS Simulator.Live, against
openjev serve --backend mlx:longStreamIsWhole,test_think,test_chatandtest_chat_streamamong them, and skipped 4 (the encoder models).tests/test_live.pypassed 12 and skipped 4.Departures (D-059)
confidence-thresholdat 0.9, not the checkpoint's entropy-bound one (ported and tested, unused).t_min,t_max,confidence_thresholdandentropy_boundare now Doubles.diffusion_full_canvas, canvas overrides, compile, chunked prefill,decoder_input_ids, images in generation, and temperature above 0.thinkoff.ChatCompletionStream.bufferCapacity(64) becomesminimumCapacitypluscapacity(blockLength:)and the instance'scapacity, andTextGeneratorgainsblockLength. Both are unreleased API from Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) #136.Notes
origin/mainwas merged in rather than rebased: the branch already holds a merge of #136, and a rebase would replay every commit through #136's conflicts.