Skip to content

[Feature] Add a Rust frontend for omni-modal, prefill-only Jev models #1

Description

@hsliuustc0106

Build a minimal, model-independent Rust frontend for Jev-style models that accept text, image, audio, video, and mixed-modality inputs and return typed decisions.

The first milestone covers the multimodal serving interface and request transport. Laya is the first backend for validating text inference. Actual modality support depends on the selected backend.

The initial implementation should:

  • Use Axum, Tokio, and Reqwest in a single Rust binary.
  • Expose POST /v1/systemone using the model, state, and questions envelope, with choice, score, and noul outputs. Jev API reference
  • Forward multimodal request bodies unchanged, including media references, inline data, and mixed-input structures. Leave media loading, preprocessing, tokenization, and inference to the backend.
  • Forward authorization when supplied and preserve backend response status, content type, and body.
  • Expose GET /health backed by the model worker’s health endpoint.
  • Reuse one HTTP client with a 60-second timeout and no automatic retries. Return 502 for connection failures and 504 for timeouts.
  • Configure the frontend through OMNI_JEV_BIND and OMNI_JEV_BACKEND_URL, defaulting to 127.0.0.1:8080 and http://127.0.0.1:8000.
  • Start the frontend and backend separately.

Acceptance criteria:

  • Mock-backend tests verify unchanged forwarding of text, image, audio, video, and mixed-input payloads.
  • Real Laya text requests cover all three decision types and match direct backend outputs.
  • The public interface and configuration contain no Laya-specific assumptions.
  • Backend errors, unavailable workers, timeouts, and health checks are tested.
  • Formatting, Clippy, Cargo tests, and a CPU smoke test pass.
  • Documentation distinguishes multimodal transport support from each backend’s inference capabilities.

Custom scheduling, Rust-side preprocessing, and additional model integrations are follow-up work. vLLM RFC #19038 provides a reference for future prefill scheduling and memory optimizations.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions