Repository navigation
Say the mlx backend serves chat and OpenJevCore has the encoder contract, and catch the configuration reference up with JevK5 - #147
Merged
Conversation
…act, and catch the configuration reference up with JevK5 Configuration.md's OPENJEV_GEN_MAX_INFLIGHT row said the chat route would serve once the model generates text (issue #51); PR #137 (D-059) made the mlx backend serve it. docs/10's encoder section said OpenJevCore will mirror upstream's EncoderEngine; QuestionReadBackend and EncoderDecisionEngine have done so since PR #82 (#67). Its batching bullet said "per forward pass", where the engine bounds each backend call and a backend may split a call into passes. Three Configuration.md sentences still described the port before PR #122 added JevK5: "Two are this port's own" (OPENJEV_JEVK5_MODEL makes three), "as upstream's do" (upstream's JevK5 keeps OPENJEV_JEVK5_WORKERS reads in flight), and JevK5's upstream variables listed under "the backends this port does not have". Docs only; the table's variable and default cells are unchanged.
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The documentation changes accurately match the implemented server routes and backend contracts.
Review effort: Balanced
Findings: None
What changed in this PR
Updates documentation to reflect the implemented chat route, encoder contract, and JevK5 backend behavior.
Changes:
- Documents MLX chat serving and current encoder concurrency.
- Replaces future-tense encoder plans with implemented contracts.
- Clarifies JevK5 configuration and batching behavior.
| File | Description |
|---|---|
Sources/OpenJevServer/Documentation.docc/Configuration.md |
Updates backend and configuration semantics. |
docs/10-other-models.md |
Documents the implemented encoder architecture and batching. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the sentences in
Sources/OpenJevServer/Documentation.docc/Configuration.mdanddocs/10-other-models.mdthat describe merged work as future work, and three more in Configuration.md that still describe the port as it was before JevK5 (#122). Docs only: no Swift code changes.PR #144 (the CLM deferral) is open, so the CLM parts of both files are left as they are:
Configuration.md:18("nor doesclmyet (issue #59)") and docs/10's CLM section. This branch and #144 merge without conflicts (git merge-treeagainst #144's head).Future tense about merged work
Configuration.md:48(OPENJEV_GEN_MAX_INFLIGHT, Meaning)mlxbackend serves the route because its model generates text (D-059); the encoder backends have no chat route.extension DiffusionGemmaRuntime: TextGenerator(Sources/OpenJevDiffusionGemma/Runtime/Generation.swift:168).DecisionEngine.textGeneratorreturns the backend when it conforms (Sources/OpenJevCore/Engine/SystemOneService.swift:175); an encoder engine keeps the protocol'snil(line 36).OpenJevApplication.routerbuildsChatCompletionsonly for a text generator and only then registersPOST /v1/chat/completions(Sources/OpenJevServer/OpenJevApplication.swift:35, 66). D-059 item 9 (docs/06-decisions.md:3988).docs/10-other-models.md:56(intro of "What a Swift encoder backend shares")EncoderEnginecontract, which the SwiftOpenJevCorewill mirror as a sibling ofDecisionBackend:EncoderEnginecontract, whichOpenJevCoremirrors asQuestionReadBackend, a sibling ofDecisionBackend, and theEncoderDecisionEnginethat reads through it (#67):public protocol QuestionReadBackend, "the sibling ofDecisionBackendthat decision D-005 names" (Sources/OpenJevCore/Engine/QuestionReadBackend.swift:24, 28), andpublic actor EncoderDecisionEngine, which reads questions in batches through it (Sources/OpenJevCore/Engine/EncoderDecisionEngine.swift:5, 16). Both came with PR #82 (be330d3, merged 2026-09-30, closing #67); the sentence dates from the planning commit 51448d9 (2026-09-29).The
OPENJEV_GEN_MAX_INFLIGHTrow keeps its first two cells, whichConfigurationReferenceTestsreads, and its first sentence, which still describes how the server uses the variable:ServerSettings(environment:)reads it with default 8 (ServerSettings.swift:232) and refuses a value below 1 at startup (line 346);ChatCompletionsConfiguration(settings)hands it toChatCompletions(ChatCompletionsRoute.swift:163to 166), whoseGenerationCapacitymakes it the permits of theslotssemaphore and addsOPENJEV_GEN_MAX_QUEUEto it for the 529 bound (Sources/OpenJevCore/Generation/ChatCompletions.swift:232to 236). A request is counted in before its prompt renders (line 123) and waits for a slot before it generates, whole or streamed (lines 178 and 209). The runtime generates a whole reply inside its actor (Generation.swift:56), so above 1 the requests holding slots wait their turn.docs/deployment.mdalready says this (its settings row and "Text generation").The encoder section, bullet by bullet
build_schemawith the same forced answers and limits (max_choices24 for Verdict, 255 otherwise)EncoderQuestionSchemaBuilder(maxChoices: backend.maxChoices)(EncoderDecisionEngine.swift:48) keepsQuestionSchemaBuilder's rules, limits, forced answers and messages (Sources/OpenJevCore/Schema/EncoderQuestionSchema.swift:62to 68).maxChoicesis 24 inVerdictBackend.swift:94and 255 inLayaBackend.swift:151andJevK5Backend.swift:99.images,steps > 1,samples > 1,thinkandsequential("{model} does not support {field}")UnsupportedOptions.checkagainst.readsOnly(EncoderDecisionEngine.swift:76), message atRequestAdmission.swift:33.OPENJEV_ENCODER_BATCH(16) questions per forward passdocs/10-other-models.md:63readhandsreadBatchat mostbatchSizequestions per call (EncoderDecisionEngine.swift:125to 156), andbatchSizeisOPENJEV_ENCODER_BATCH(BackendProvider.swift:91). Verdict splits a call into Core ML calls of at mostmaxBatchRows, 16 on macOS and 1 on iOS (VerdictBackend.swift:70to 76, 141 to 145); JevK5 reads each question in passes of its own (JevK5Backend.swift:172, D-052 item 6). Now: "per backend call, which a backend may run as several model passes".[P(true), 1 − P(true)]BatchReadResult.probabilities(QuestionReadBackend.swift:7to 10), checked byvalidate(EncoderDecisionEngine.swift:163).EncoderDecisionEngine.swift:65to 66./v1/modelsentry with the upstream description text and release date, and acceptance ofjev-latestandjev-previewwhen it is the only model in the processEncoderDecisionEngine.servedModelsisServedModels.encoder(backend.modelInfo)(SystemOneService.swift:184, 92): its own name plus the SDK aliases, and a listing of itself alone withKnownEncoderModels' upstream texts. A process loads one backend, so the condition always holds, as in upstream'sserved_models.Also stale since #122 (JevK5), not future tense
Each was true when PR #107 wrote it (37eb2be, 2026-10-02) and stopped being true when PR #122 added the
jevk5backend (5b4a70c, 2026-10-03). They fall outside the future-tense sweep, but they deny merged work in the same way, and the code settles each.Configuration.md:9OPENJEV_ENCODER_FUNCTIONS(D-042),OPENJEV_ENCODER_MODELS(D-033) and, since #122,OPENJEV_JEVK5_MODEL(D-052;ServerSettings.swift:96). They are the only table variables that upstream'sopenjev/*.pyat dcd2094 never reads.Configuration.md:36(OPENJEV_MAX_INFLIGHT, Meaning)maxInflight1:EncoderEngineConfiguration(settings)never sets it (BackendProvider.swift:86to 93; the default is inEncoderEngineConfiguration.swift:25). Upstream'sEncoderEngine.workersis 1 (encoders.py:61) and its Laya and Verdict keep it, butJevK5Enginesetsworkers = settings.jevk5_workers,OPENJEV_JEVK5_WORKERS, default 32 (encoders.py:364,config.py:87).Configuration.md:88(Other variables)OPENJEV_MODELandOPENJEV_JEVK5_WORKERS(JevK5).OPENJEV_MODELandOPENJEV_JEVK5_WORKERS(JevK5, which this port runs on MLX in the process, D-052).jevk5backend (Sources/openjev/BackendRegistry.swift:134), run in the process (D-052 item 6,docs/06-decisions.md:2956). Upstream's JevK5 readsOPENJEV_MODELas "the weights the vLLM server at OPENJEV_UPSTREAM serves" (config.py:84to 86), and its CLM reads embeddings from the same server (encoders.py:335). The list of variables, the CLM settings among them, is unchanged, and #144 does not touch this paragraph. The rest of the paragraph is rewrapped, not reworded.What I read and left
I split both files into sentences, with wrapped lines joined and table rows kept whole (65 units in Configuration.md, 51 in docs/10), flagged future tense, "yet", "until", "once", "not yet" and similar, and read every unit. Left as written:
Configuration.md:18and docs/10's CLM section (docs/10-other-models.md:44to 52): held for Defer CLM until someone asks for it: a D-011 addendum, and the pages that said it was coming (#59) #144.docs/10-other-models.md:42, "Builds for iOS; not yet run on an iPhone.": still true. README.md,docs/compatibility.md:165anddocs/development.md:46say the same, and no issue or PR records a JevK5 run on an iPhone.docs/10-other-models.md:29, "refused withEncoderLoadError.noPackageuntil then": "then" is the app'sprefetch(lengths:)call, a runtime condition.docs/10-other-models.md:3, "calibration that a Swift port must reproduce exactly": a requirement, which still holds.Checks
make lintpasses.swift testnot run, since this is docs only. CI runsConfigurationReferenceTestsbecause Configuration.md changed: the variable and default cells of all 33 rows are identical to main's (a Python copy of the test's row parser), and every row still has four cells.Sources/; no symbol link was added or changed.