You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Build a minimal, model-independent Rust frontend for Jev-style models that accept text, image, audio, video, and mixed-modality inputs and return typed decisions.
The first milestone covers the multimodal serving interface and request transport. Laya is the first backend for validating text inference. Actual modality support depends on the selected backend.
The initial implementation should:
Use Axum, Tokio, and Reqwest in a single Rust binary.
Expose POST /v1/systemone using the model, state, and questions envelope, with choice, score, and noul outputs. Jev API reference
Forward multimodal request bodies unchanged, including media references, inline data, and mixed-input structures. Leave media loading, preprocessing, tokenization, and inference to the backend.
Forward authorization when supplied and preserve backend response status, content type, and body.
Expose GET /health backed by the model worker’s health endpoint.
Reuse one HTTP client with a 60-second timeout and no automatic retries. Return 502 for connection failures and 504 for timeouts.
Configure the frontend through OMNI_JEV_BIND and OMNI_JEV_BACKEND_URL, defaulting to 127.0.0.1:8080 and http://127.0.0.1:8000.
Start the frontend and backend separately.
Acceptance criteria:
Mock-backend tests verify unchanged forwarding of text, image, audio, video, and mixed-input payloads.
Real Laya text requests cover all three decision types and match direct backend outputs.
The public interface and configuration contain no Laya-specific assumptions.
Backend errors, unavailable workers, timeouts, and health checks are tested.
Formatting, Clippy, Cargo tests, and a CPU smoke test pass.
Documentation distinguishes multimodal transport support from each backend’s inference capabilities.
Custom scheduling, Rust-side preprocessing, and additional model integrations are follow-up work. vLLM RFC #19038 provides a reference for future prefill scheduling and memory optimizations.
Build a minimal, model-independent Rust frontend for Jev-style models that accept text, image, audio, video, and mixed-modality inputs and return typed decisions.
The first milestone covers the multimodal serving interface and request transport. Laya is the first backend for validating text inference. Actual modality support depends on the selected backend.
The initial implementation should:
POST /v1/systemoneusing themodel,state, andquestionsenvelope, withchoice,score, andnouloutputs. Jev API referenceGET /healthbacked by the model worker’s health endpoint.502for connection failures and504for timeouts.OMNI_JEV_BINDandOMNI_JEV_BACKEND_URL, defaulting to127.0.0.1:8080andhttp://127.0.0.1:8000.Acceptance criteria:
Custom scheduling, Rust-side preprocessing, and additional model integrations are follow-up work. vLLM RFC #19038 provides a reference for future prefill scheduling and memory optimizations.