Skip to content

Say the mlx backend serves chat and OpenJevCore has the encoder contract, and catch the configuration reference up with JevK5 - #147

Merged
alaineid merged 1 commit into
mainfrom
docs-stale-generation
Oct 6, 2026
Merged

alaineid merged 1 commit into
mainfrom
docs-stale-generation

Conversation

@alaineid

@alaineid alaineid commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Fixes the sentences in Sources/OpenJevServer/Documentation.docc/Configuration.md and docs/10-other-models.md that describe merged work as future work, and three more in Configuration.md that still describe the port as it was before JevK5 (#122). Docs only: no Swift code changes.

PR #144 (the CLM deferral) is open, so the CLM parts of both files are left as they are: Configuration.md:18 ("nor does clm yet (issue #59)") and docs/10's CLM section. This branch and #144 merge without conflicts (git merge-tree against #144's head).

Future tense about merged work

Line Said Says now Evidence
Configuration.md:48 (OPENJEV_GEN_MAX_INFLIGHT, Meaning) The route serves once the model generates text (issue #51); until then this is read and checked. The mlx backend serves the route because its model generates text (D-059); the encoder backends have no chat route. Issue #51 closed and PR #137 merged on 2026-10-06 (bbab174). extension DiffusionGemmaRuntime: TextGenerator (Sources/OpenJevDiffusionGemma/Runtime/Generation.swift:168). DecisionEngine.textGenerator returns the backend when it conforms (Sources/OpenJevCore/Engine/SystemOneService.swift:175); an encoder engine keeps the protocol's nil (line 36). OpenJevApplication.router builds ChatCompletions only for a text generator and only then registers POST /v1/chat/completions (Sources/OpenJevServer/OpenJevApplication.swift:35, 66). D-059 item 9 (docs/06-decisions.md:3988).
docs/10-other-models.md:56 (intro of "What a Swift encoder backend shares") Upstream's EncoderEngine contract, which the Swift OpenJevCore will mirror as a sibling of DecisionBackend: Upstream's EncoderEngine contract, which OpenJevCore mirrors as QuestionReadBackend, a sibling of DecisionBackend, and the EncoderDecisionEngine that reads through it (#67): public protocol QuestionReadBackend, "the sibling of DecisionBackend that decision D-005 names" (Sources/OpenJevCore/Engine/QuestionReadBackend.swift:24, 28), and public actor EncoderDecisionEngine, which reads questions in batches through it (Sources/OpenJevCore/Engine/EncoderDecisionEngine.swift:5, 16). Both came with PR #82 (be330d3, merged 2026-09-30, closing #67); the sentence dates from the planning commit 51448d9 (2026-09-29).

The OPENJEV_GEN_MAX_INFLIGHT row keeps its first two cells, which ConfigurationReferenceTests reads, and its first sentence, which still describes how the server uses the variable: ServerSettings(environment:) reads it with default 8 (ServerSettings.swift:232) and refuses a value below 1 at startup (line 346); ChatCompletionsConfiguration(settings) hands it to ChatCompletions (ChatCompletionsRoute.swift:163 to 166), whose GenerationCapacity makes it the permits of the slots semaphore and adds OPENJEV_GEN_MAX_QUEUE to it for the 529 bound (Sources/OpenJevCore/Generation/ChatCompletions.swift:232 to 236). A request is counted in before its prompt renders (line 123) and waits for a slot before it generates, whole or streamed (lines 178 and 209). The runtime generates a whole reply inside its actor (Generation.swift:56), so above 1 the requests holding slots wait their turn. docs/deployment.md already says this (its settings row and "Text generation").

The encoder section, bullet by bullet

Bullet Result Code
build_schema with the same forced answers and limits (max_choices 24 for Verdict, 255 otherwise) Holds EncoderQuestionSchemaBuilder(maxChoices: backend.maxChoices) (EncoderDecisionEngine.swift:48) keeps QuestionSchemaBuilder's rules, limits, forced answers and messages (Sources/OpenJevCore/Schema/EncoderQuestionSchema.swift:62 to 68). maxChoices is 24 in VerdictBackend.swift:94 and 255 in LayaBackend.swift:151 and JevK5Backend.swift:99.
A 400 for images, steps > 1, samples > 1, think and sequential ("{model} does not support {field}") Holds UnsupportedOptions.check against .readsOnly (EncoderDecisionEngine.swift:76), message at RequestAdmission.swift:33.
Batched reads of at most OPENJEV_ENCODER_BATCH (16) questions per forward pass Fixed, docs/10-other-models.md:63 The engine bounds each backend call, not each pass: read hands readBatch at most batchSize questions per call (EncoderDecisionEngine.swift:125 to 156), and batchSize is OPENJEV_ENCODER_BATCH (BackendProvider.swift:91). Verdict splits a call into Core ML calls of at most maxBatchRows, 16 on macOS and 1 on iOS (VerdictBackend.swift:70 to 76, 141 to 145); JevK5 reads each question in passes of its own (JevK5Backend.swift:172, D-052 item 6). Now: "per backend call, which a backend may run as several model passes".
A distribution over the caller's options in the caller's order; noul is [P(true), 1 − P(true)] Holds BatchReadResult.probabilities (QuestionReadBackend.swift:7 to 10), checked by validate (EncoderDecisionEngine.swift:163).
Deterministic: the seed is unused Holds EncoderDecisionEngine.swift:65 to 66.
Its own /v1/models entry with the upstream description text and release date, and acceptance of jev-latest and jev-preview when it is the only model in the process Holds EncoderDecisionEngine.servedModels is ServedModels.encoder(backend.modelInfo) (SystemOneService.swift:184, 92): its own name plus the SDK aliases, and a listing of itself alone with KnownEncoderModels' upstream texts. A process loads one backend, so the condition always holds, as in upstream's served_models.

Also stale since #122 (JevK5), not future tense

Each was true when PR #107 wrote it (37eb2be, 2026-10-02) and stopped being true when PR #122 added the jevk5 backend (5b4a70c, 2026-10-03). They fall outside the future-tense sweep, but they deny merged work in the same way, and the code settles each.

Line Said Says now Evidence
Configuration.md:9 Two are this port's own. Three are this port's own. The table marks three variables "This port's": OPENJEV_ENCODER_FUNCTIONS (D-042), OPENJEV_ENCODER_MODELS (D-033) and, since #122, OPENJEV_JEVK5_MODEL (D-052; ServerSettings.swift:96). They are the only table variables that upstream's openjev/*.py at dcd2094 never reads.
Configuration.md:36 (OPENJEV_MAX_INFLIGHT, Meaning) The encoders make one call at a time whatever it says, as upstream's do. The encoders make one call at a time whatever it says, as upstream's Verdict and Laya do. Here every encoder engine keeps maxInflight 1: EncoderEngineConfiguration(settings) never sets it (BackendProvider.swift:86 to 93; the default is in EncoderEngineConfiguration.swift:25). Upstream's EncoderEngine.workers is 1 (encoders.py:61) and its Laya and Verdict keep it, but JevK5Engine sets workers = settings.jevk5_workers, OPENJEV_JEVK5_WORKERS, default 32 (encoders.py:364, config.py:87).
Configuration.md:88 (Other variables) Upstream also reads variables for the backends this port does not have, which it ignores: ..., and OPENJEV_MODEL and OPENJEV_JEVK5_WORKERS (JevK5). Upstream also reads variables for its vLLM server and the backends that read through it, which this port ignores: ..., and OPENJEV_MODEL and OPENJEV_JEVK5_WORKERS (JevK5, which this port runs on MLX in the process, D-052). The port has a jevk5 backend (Sources/openjev/BackendRegistry.swift:134), run in the process (D-052 item 6, docs/06-decisions.md:2956). Upstream's JevK5 reads OPENJEV_MODEL as "the weights the vLLM server at OPENJEV_UPSTREAM serves" (config.py:84 to 86), and its CLM reads embeddings from the same server (encoders.py:335). The list of variables, the CLM settings among them, is unchanged, and #144 does not touch this paragraph. The rest of the paragraph is rewrapped, not reworded.

What I read and left

I split both files into sentences, with wrapped lines joined and table rows kept whole (65 units in Configuration.md, 51 in docs/10), flagged future tense, "yet", "until", "once", "not yet" and similar, and read every unit. Left as written:

  • Configuration.md:18 and docs/10's CLM section (docs/10-other-models.md:44 to 52): held for Defer CLM until someone asks for it: a D-011 addendum, and the pages that said it was coming (#59) #144.
  • docs/10-other-models.md:42, "Builds for iOS; not yet run on an iPhone.": still true. README.md, docs/compatibility.md:165 and docs/development.md:46 say the same, and no issue or PR records a JevK5 run on an iPhone.
  • docs/10-other-models.md:29, "refused with EncoderLoadError.noPackage until then": "then" is the app's prefetch(lengths:) call, a runtime condition.
  • docs/10-other-models.md:3, "calibration that a Swift port must reproduce exactly": a requirement, which still holds.

Checks

  • make lint passes.
  • swift test not run, since this is docs only. CI runs ConfigurationReferenceTests because Configuration.md changed: the variable and default cells of all 33 rows are identical to main's (a Python copy of the test's row parser), and every row still has four cells.
  • The Documentation workflow builds the DocC site, as Configuration.md is under Sources/; no symbol link was added or changed.
  • No em dashes.

…act, and catch the configuration reference up with JevK5

Configuration.md's OPENJEV_GEN_MAX_INFLIGHT row said the chat route would serve once the model
generates text (issue #51); PR #137 (D-059) made the mlx backend serve it. docs/10's encoder
section said OpenJevCore will mirror upstream's EncoderEngine; QuestionReadBackend and
EncoderDecisionEngine have done so since PR #82 (#67). Its batching bullet said "per forward
pass", where the engine bounds each backend call and a backend may split a call into passes.

Three Configuration.md sentences still described the port before PR #122 added JevK5: "Two are
this port's own" (OPENJEV_JEVK5_MODEL makes three), "as upstream's do" (upstream's JevK5 keeps
OPENJEV_JEVK5_WORKERS reads in flight), and JevK5's upstream variables listed under "the backends
this port does not have". Docs only; the table's variable and default cells are unchanged.
Copilot AI balanced review requested due to automatic review settings October 6, 2026 18:48

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The documentation changes accurately match the implemented server routes and backend contracts.

Review effort: Balanced
Findings: None

What changed in this PR

Updates documentation to reflect the implemented chat route, encoder contract, and JevK5 backend behavior.

Changes:

  • Documents MLX chat serving and current encoder concurrency.
  • Replaces future-tense encoder plans with implemented contracts.
  • Clarifies JevK5 configuration and batching behavior.
File Description
Sources/​OpenJevServer/​Documentation.docc/​Configuration.md Updates backend and configuration semantics.
docs/​10-other-models.md Documents the implemented encoder architecture and batching.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@alaineid
alaineid merged commit d176ce3 into main Oct 6, 2026
5 of 7 checks passed
@alaineid
alaineid deleted the docs-stale-generation branch October 6, 2026 19:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants