Skip to content

Generate text, think and chat on the MLX backend: the diffusion sampler, the block loop, think and the chat model (#50, #51, #52, #53, D-059) - #137

Merged
alaineid merged 11 commits into
mainfrom
issue-50-51-52-generation-and-think
Oct 6, 2026
Merged

alaineid merged 11 commits into
mainfrom
issue-50-51-52-generation-and-think

Conversation

@alaineid

@alaineid alaineid commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Closes #50.
Closes #51.
Closes #52.
Closes #53.

Milestone 5's model half (#49): text generation on the MLX backend, the think option, and the chat endpoint (#136) wired to the model. Decision D-059; D-058 is the chat route's, from #136.

What was built

  • Oracle. Tools/fixtures/generation_oracle.py writes Fixtures/generation/generation.json from upstream's own path (MlxRuntime.generate, MlxEngine.decide) on mlx-vlm 0.6.15. It holds:

    • test vectors for the sampling functions of diffusion.py lines 285 to 505, computed on the CPU;
    • eight seeded greedy chat replies, recorded block by block (initial canvas, passes, ending, final canvas) with every emit call: a short answer, a list, JSON, a 640-token story over blocks of 256/256/128, a reply ended by an extra stop id, one cut by max_tokens 40, a 320-token report after a 1,235-token prompt, and one under another seed. The report's first commit trims the sliding layers past their 1,023-position window;
    • two think requests (the README example with think 128, and 24 sequential nouls with think 64), with every answer after the thought.

    Everything runs twice, the second time reversed, and nothing is written unless both runs agree bit for bit. Dashes the model writes are stored as JSON escapes. --check reproduced the committed file, and the committed run record holds that check.

  • Seeding. Upstream never seeds MLX's generator, so its replies vary from process to process. The port gives each reply MLXRandom.RandomState(seed: Configuration.generationSeed) (0 by default), which draws what mx.random.seed gives. That is tested on the sampler and on every block's recorded initial canvas.

  • Diffusion sampler: entropy-bound canvas update, temperature schedule, stable-and-confident stopping #50. DiffusionSampler, held to mlx-vlm's vectors bit for bit.

  • Block generation loop: canvas sizing, cache commits, streaming detokenizer, stop ids #51. denoiseBlock, updateCache (diffusion_update_cache, on copies that never share array objects), the SentencePiece streaming detokenizer, and DiffusionGemmaRuntime.generate. Canvases are min(256, max(remaining, 64)); it stops at EOS {1, 106, 50}, stop ids and max_tokens; it calls emit per token plus a tail; a cancelled task ends the reply before the prefill, before the next block or after the last one; and it reuses a read's cached prefill.

  • The think option: thought generation before a read, prefix continuation, billing #52. think runs on generate, and capabilities.think is on.

  • /v1/chat/completions: OpenAI-compatible generation with streaming, JSON mode and cancellation #53. DiffusionGemmaRuntime conforms to Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) #136's TextGenerator, so the server serves POST /v1/chat/completions on mlx.

Review fixes in this round

  1. Streamed replies stopped after 64 pieces.
    • generate emits a committed block of up to 256 tokens back to back, and the chat stream's queue was upstream's 64. So the 65th piece of every block of more than 64 tokens ended the reply without its finish chunk or [DONE].
    • The queue now holds max(64, 2 × blockLength + 1) pieces, 513 here, from a new TextGenerator.blockLength that the generator reports (256 for the runtime). I chose a reported size over a constant because the block length is the model's: a token-by-token generator reports 1 and keeps 64.
    • wholeBlockStreams emits 257 pieces from one synchronous call; it failed with readerFellBehind before the change and passes now. The slow-reader test holds at the new size.
    • Checked live with curl -N: a 451-token story streams to its finish and [DONE], equal to the same request without stream. The new live test longStreamIsWhole checks this.
    • Upstream's queue of 64 has the same flaw (D-059 item 11).
  2. Chat skips only the two channel markers, 100 and 101 (maintainer's call).
    • Upstream's list [100, 45518, 107, 101] also drops the single newline token 107 and thought 45518 wherever they appear, so lists run together and {"thought": 1} becomes {"": 1}.
    • Only 107 goes; other newline tokens stay. The recorded story keeps 20 newlines, from 10 tokens of id 108.
    • No reply on the checkpoint showed a thought channel's text with the shorter list. generate's parity tests keep upstream's list (D-059 item 10).
  3. Docs and tests.
    • Docs updated: DocC and README no longer say think or chat are missing, and the newline wording is exact.
    • The exact-tier think test now compares every answer after the thought, team and tone included, probabilities and confidence: all bit for bit.
    • commitsCopy compares the prefill's tensor digests after two continuations, short prompt and long.
    • The think test now detects skipping.
    • The doc slips are fixed: the schedule's floor, the defaults, the cancellation wording, and the unconditional conformance.
  4. Em dashes. The two U+2014 characters in the fixture are now escapes; the oracle escapes U+2012 to U+2015 as mlx_vlm_oracle.py does.

Measured (2026-10-06, the reference Mac, OPENJEV_MLX_CACHE_LIMIT_GB=4)

Exact tier (OPENJEV_MLX_METALLIB = the wheel's metallib, oracle RoPE table):

  • All 8 replies agree whole: initial canvases, final canvases, passes, endings, finish reasons and every emit call.
  • Both thoughts agree token for token (128 and 64 ids), with upstream's billing (469 and 1,896 input tokens; 128 and 64 output) and every answer after them, bit for bit.
  • ReadOracleTests (the 63 reads, bit for bit) and SamplerTests pass.

Native tier:

Reply Tokens Blocks Agreement with the oracle
short answer 7 1 whole reply
list 10 1 whole reply
JSON 19 1 whole reply
story 640 3 first 8 tokens
stop id 1 1 whole reply
cut by max_tokens 40 1 first 4 tokens
long prompt 320 2 first 2 tokens
short answer, seed 7 7 1 whole reply

The long replies part where a near-tied argmax flips under mlx-swift's kernels. #51's long-reply agreement is therefore held in the exact tier, as D-014 holds reads (D-059 item 12).

Natively the seeded initial canvases match wherever the blocks before them drew alike. Both thoughts bill as upstream does (128/469 and 64/1,896), though their ids part after 48 and 3 tokens.

#52's "available for logging at debug level": the thought's ids are available to the engine, in the read's prefix, never in the Decision. OpenJevCore has no logger, so nothing logs them.

Whole package, natively with the model tests on, run in groups because Xcode crashes on a single full run:

Group Passed Skipped Failed Not run
OpenJevCoreTests 383 0 0 0
OpenJevServerTests 190 0 0 0
CLI, bench, JevK5, encoders 199 3 0 1
DiffusionGemma, model-free 163 1 0 0
DiffusionGemma MLX and oracle 63 1 0 0
DiffusionGemma live (runtime, policy, regression, images) 18 0 0 0
Upstream's read, think and chat cases 15 0 0 0
Live-suite target, no server 33 15 0 0
  • The skips are JevK5 live (OPENJEV_JEVK5_MODEL), the Hub download (OPENJEV_TEST_DOWNLOAD), the vision full tensors, and the live suite without OPENJEV_LIVE_URL.
  • The test not run is encodersWithoutCoreML, which compiles only without Core ML.
  • .github/scripts/check-test-log.sh passes on all eight groups' logs.
  • make lint is clean.
  • OpenJevCore-iOS and OpenJevDiffusionGemma build for the iOS Simulator.

Live, against openjev serve --backend mlx:

  • The Swift live suite passed 13, longStreamIsWhole, test_think, test_chat and test_chat_stream among them, and skipped 4 (the encoder models).
  • Upstream's tests/test_live.py passed 12 and skipped 4.
  • No failures in either.

Departures (D-059)

  • Sampler. The sampler upstream runs: confidence-threshold at 0.9, not the checkpoint's entropy-bound one (ported and tested, unused).
  • Temperature. The temperature schedule divides logits even at temperature 0, computed in Double, so the generation configuration's t_min, t_max, confidence_threshold and entropy_bound are now Doubles.
  • Options left out. These are options upstream never sets: the static cache, diffusion_full_canvas, canvas overrides, compile, chunked prefill, decoder_input_ids, images in generation, and temperature above 0.
  • Seeding. Every reply is seeded with 0.
  • Prefill sharing. Generation reuses a read's cached prefill and caches none of its own.
  • Unsupported sampler. A checkpoint whose sampler class mlx-vlm refuses loads and reads, with think off.
  • Chat skip list. Chat skips only the two channel markers (item 10).
  • Stream queue. A streamed reply's queue holds two blocks and the tail (item 11).
  • Long-reply agreement. Held in the exact tier (item 12).
  • Public API. ChatCompletionStream.bufferCapacity (64) becomes minimumCapacity plus capacity(blockLength:) and the instance's capacity, and TextGenerator gains blockLength. Both are unreleased API from Serve /v1/chat/completions over a TextGenerator: the HTTP and core side of #53 (D-058) #136.

Notes

origin/main was merged in rather than rebased: the branch already holds a merge of #136, and a rebase would replay every commit through #136's conflicts.

…e block loop and the think option (#50, #51, #52, D-058)

Port mlx-vlm 0.6.15's sampling functions (DiffusionSampler), one block of
stream_diffusion_generate as upstream's MlxRuntime.generate runs it
(greedy, the default confidence-threshold sampler, the linear temperature
schedule, self-conditioning, stable-and-confident stopping), and
diffusion_update_cache on copies of the layer caches. DiffusionGemmaRuntime
gains generate(prompt:maxTokens:stopIDs:skipSpecialTokenIDs:emit:) with
canvas sizing, EOS and stop ids, the SentencePiece streaming detokenizer,
emit and task cancellation, and think(prompt:budget:stopIDs:) on it, so
capabilities.think is on. Each reply's canvases come from a RandomState
seeded with Configuration.generationSeed (0), which upstream leaves
unseeded.

Tools/fixtures/generation_oracle.py records the oracle in
Fixtures/generation: sampler vectors, seven seeded greedy replies block by
block, and two think requests. On the oracle's Metal library every reply
and thought matches token for token with upstream's billing; natively the
short replies match whole and the long story for its first 8 tokens.
Upstream's three think cases and the live suite's test_think now run.
Copilot AI balanced review requested due to automatic review settings October 6, 2026 14:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Cache-copy isolation and cancellation handling have unresolved correctness issues.

Review effort: Balanced
Findings: 1 High severity · 3 Medium severity

Open (4)
What changed in this PR

Adds MLX text generation and enables think, providing the model-side foundation for the future chat endpoint.

Changes:

  • Implements sampling, block denoising, cache continuation and streaming text.
  • Enables seeded thought generation through the runtime.
  • Adds upstream oracle comparisons, tests and documentation.
File Description
Tools/​README.md Lists the generation oracle tool.
Tools/​oracle/​results/​generation_run.json Records oracle timings and reproducibility.
Tools/​fixtures/​generation_oracle.py Generates sampling, reply and thought fixtures.
Tests/​OpenJevLiveTests/​LiveTests.swift Enables live think testing.
Tests/​OpenJevDiffusionGemmaTests/​Support/​GenerationFixtures.swift Decodes generation fixtures.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​UpstreamReadCaseTests.swift Enables upstream think cases.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​RuntimeTests.swift Updates capability and refusal tests.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​RuntimeStubs.swift Adds scripted generation support.
Tests/​OpenJevDiffusionGemmaTests/​Runtime/​ImageRuntimeTests.swift Checks image/think refusals.
Tests/​OpenJevDiffusionGemmaTests/​Generation/​StreamingDetokenizerTests.swift Tests streamed text and byte handling.
Tests/​OpenJevDiffusionGemmaTests/​Generation/​SamplerTests.swift Checks sampling against oracle vectors.
Tests/​OpenJevDiffusionGemmaTests/​Generation/​GenerationRuntimeTests.swift Tests block-loop behavior without weights.
Tests/​OpenJevDiffusionGemmaTests/​Generation/​GenerationOracleTests.swift Compares model replies and thoughts upstream.
Sources/​OpenJevDiffusionGemma/​Runtime/​RuntimeConfiguration.swift Adds the generation seed.
Sources/​OpenJevDiffusionGemma/​Runtime/​Generation.swift Implements generation and think.
Sources/​OpenJevDiffusionGemma/​Runtime/​DiffusionGemmaRuntime.swift Wires generation and capabilities.
Sources/​OpenJevDiffusionGemma/​Model/​Prefill.swift Adds committed-block cache updates.
Sources/​OpenJevDiffusionGemma/​Model/​ModelTree.swift Supplies sliding-cache window sizes.
Sources/​OpenJevDiffusionGemma/​Model/​LayerCache.swift Supports cache copying and extension.
Sources/​OpenJevDiffusionGemma/​Model/​Configuration.swift Uses Double configuration values.
Sources/​OpenJevDiffusionGemma/​Generation/​StreamingDetokenizer.swift Converts tokens into streamed text.
Sources/​OpenJevDiffusionGemma/​Generation/​DiffusionSampler.swift Implements upstream sampling functions.
Sources/​OpenJevDiffusionGemma/​Generation/​BlockDenoising.swift Adds generation policy and block denoising.
Sources/​OpenJevDiffusionGemma/​Documentation.docc/​ReadingWithDiffusionGemma.md Documents generation usage.
Sources/​OpenJevDiffusionGemma/​Documentation.docc/​OpenJevDiffusionGemma.md Exposes generation documentation.
Fixtures/​README.md Lists generation fixtures.
Fixtures/​generation/​README.md Explains oracle contents and seeding.
docs/​compatibility.md Updates think support and differences.
docs/​09-conformance-and-testing.md Documents generation validation tiers.
docs/​07-risks-and-unknowns.md Updates generation parity risks.
docs/​06-decisions.md Records D-059 and intentional departures.
docs/​05-architecture.md Describes generation integration.
docs/​03-diffusiongemma.md Explains the upstream generation policy.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread Sources/OpenJevDiffusionGemma/Model/LayerCache.swift Outdated
Comment on lines +73 to +75
let trim = keys.dim(2) - window + 1
let keptKeys = trim > 0 ? keys[.ellipsis, trim..., 0...] : keys
let keptValues = trim > 0 ? values[.ellipsis, trim..., 0...] : values

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added in a99890b: LayerCacheTests.slidingTrim uses distinct positions and a window of 4. It updates twice past the window and checks the positions kept, the values and the offsets each time, then checks the prefill is unchanged after updating its extended() copy. The suite needs no checkpoint and runs in CI.

Comment thread Sources/OpenJevDiffusionGemma/Runtime/Generation.swift
Comment thread Sources/OpenJevDiffusionGemma/Runtime/Generation.swift
…t block, serialize the stub generation tests, fix a DocC link (#137 review, CI)

LayerCache.extended() gives the copy array objects of its own: assigning into a slice of an MLXArray changes that object, so two continuations of one prefill wrote into each other's buffer (the new forkedContinuations test fails without this). generate returns cancelled for a task cancelled before the prefill or during the last block. The stub generation suite makes MLX random state, so it now runs inside the serialized MLX suite; in parallel it took down CI's test process. DiffusionSampler's link to MLXRandom.RandomState is code text, which DocC can resolve.
…tion-and-think

# Conflicts:
#	Tests/OpenJevLiveTests/LiveTests.swift
#	docs/06-decisions.md
#	docs/09-conformance-and-testing.md
#	docs/compatibility.md
…Generator (#53, D-059)

The runtime's generate returns the core's TextGeneration, and it supplies the chat prompt (SwiftTransformersTokenizer.generationPromptIDs), Engine.enc and the thought-channel marker ids, so DecisionEngine.textGenerator returns it and the server adds the chat routes. The chat cases #136 left disabled run: test_chat_completion_on_mlx over the stub model, the four DiffusionGemma chat cases on the checkpoint, and the live suite's test_chat and test_chat_stream. D-059 item 9 answers D-058 item 10.
@alaineid alaineid changed the title Generate text and think on the MLX backend: the diffusion sampler, the block loop and think (#50, #51, #52, D-059) Generate text, think and chat on the MLX backend: the diffusion sampler, the block loop, think and the chat model (#50, #51, #52, #53, D-059) Oct 6, 2026
…skip only the channel markers in chat; regenerate the generation oracle with a long prompt and escaped dashes; tighten the generation tests (#137 review)
@alaineid
alaineid merged commit bbab174 into main Oct 6, 2026
8 checks passed
@alaineid
alaineid deleted the issue-50-51-52-generation-and-think branch October 6, 2026 17:52
alaineid added a commit that referenced this pull request Oct 6, 2026
…act, and catch the configuration reference up with JevK5 (#147)

Configuration.md's OPENJEV_GEN_MAX_INFLIGHT row said the chat route would serve once the model
generates text (issue #51); PR #137 (D-059) made the mlx backend serve it. docs/10's encoder
section said OpenJevCore will mirror upstream's EncoderEngine; QuestionReadBackend and
EncoderDecisionEngine have done so since PR #82 (#67). Its batching bullet said "per forward
pass", where the engine bounds each backend call and a backend may split a call into passes.

Three Configuration.md sentences still described the port before PR #122 added JevK5: "Two are
this port's own" (OPENJEV_JEVK5_MODEL makes three), "as upstream's do" (upstream's JevK5 keeps
OPENJEV_JEVK5_WORKERS reads in flight), and JevK5's upstream variables listed under "the backends
this port does not have". Docs only; the table's variable and default cells are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants