This document describes the implemented admin API for loading, unloading, and inspecting models at runtime.
The API avoids edits to local.json and service restarts for routine model
management.
It controls live runtime state only:
- no automatic writes back to
settings.jsonorlocal.json - no arbitrary model definitions via API
- no force unload
- no background job system for model loads
Managed backend behavior:
- this admin API is implemented and is the live control plane used by the workbench
replicasis a common runtime-only load override;target_inflightis available forllama_server,openai_remote,vllm_serve,trtllm_serve, andsglang_serve; backend-specific overrides are implemented forllama_cpp,exllamav3,vllm,vllm_serve,trtllm_serve,sglang_serve, andllama_serverllama_serverload/unload starts and stops a managed nativellama-serversubprocess; binary path, model path, library path,mmproj, draft model path, host, port, and extra native args stay in model configvllm_serveload/unload starts and stops a managed localvllm servesubprocess; binary path, target model id/path, library path, environment, host, port, API key, and extra CLI args stay in model configtrtllm_serveload/unload starts and stops a managed localtrtllm-serveprocess group; binary path, target model id/path, library path, environment, host, port, TensorRT-LLM config, parser names, and extra CLI args stay in model configsglang_serveload/unload starts and stops a managed localsglang serveprocess group; binary path, target model id/path, library path, environment, host, port, parser names, and extra CLI args stay in model config- clients use the live
load_constraintspayload to build backend-specific load controls
- Purpose
- Core Concepts
- State Semantics
- Request Behavior By Runtime State
- TranslateGemma Request Notes
- Endpoints
- Unload And In-Flight Requests
- Scheduler Alignment
- Resource Release Guarantees
- Documentation Expectations
- Out Of Scope
The current service merges settings.json and local.json into one effective config, then loads enabled models at startup.
The admin API adds a separate live control plane on top of that merged config:
- the merged config tells us which models are known to the service
- the live runtime state tells us which of those models are currently loaded
That distinction must stay explicit in both the API and the UI.
A configured model definition comes from the merged settings.json + local.json payload.
The process reads this definition at startup and when the settings reload endpoint is called. It includes fields such as:
model_pathbackenddeviceprompt_format- backend-specific settings
enabled
The admin API does not write this definition. The reload endpoint only rereads the files.
Each configured model also has a live runtime state inside the process.
Allowed states:
unloadedloadingloadedunloadingfailed
These states are runtime-only and may differ from the original enabled value in config.
- the model exists in merged config
- no runtime is currently loaded
- inference requests for this model are rejected
- the model may be loaded through the admin API
- a runtime load has started but is not complete yet
- inference requests for this model are rejected
- duplicate load requests without load overrides are idempotent and return the current state
- a runtime exists and may serve inference requests
- the model may be unloaded through the admin API
- no new inference requests are accepted for this model
- in-flight requests are allowed to finish
- once in-flight requests reach zero, runtime resources are released
- the last load attempt failed
last_erroris retained for inspection- the model may be loaded again through the admin API
For POST /v1/responses:
loaded: acceptunloaded: rejectloading: rejectunloading: rejectfailed: reject
Errors are explicit and machine-readable.
Request error codes include:
unknown_modelmodel_not_loadedmodel_loadingmodel_unloadingmodel_failedremote_execution_disallowedfile_input_unsupportedmodality_unsupportedthinking_unsupportedresponse_format_unsupportedfairness_key_queue_fullexecutor_queue_full
Requests may optionally include thinking: "default" | "enabled" | "disabled".
default preserves the model configuration. enabled and disabled are
accepted only when the selected model advertises those values in
capabilities.thinking_modes; otherwise the request is rejected with
400 thinking_unsupported.
response_format accepts strict JSON Schema output for non-streaming
vllm_serve requests. Other backends reject it with
400 response_format_unsupported. Models advertise accepted formats through
capabilities.response_formats.
llama_cpp models configured with prompt_format: "translategemma_template" use the official structured TranslateGemma request shape internally. They remain single-turn text models and continue to report capabilities.multi_turn: false.
Known-source requests should include both source_lang_code and target_lang_code.
Mixed-source requests may omit source_lang_code, or set it to "auto" or "mixed", while still providing target_lang_code. In that mode, the runtime keeps the structured TranslateGemma path, uses an internal valid source-language fallback, and prepends a short instruction asking the model to detect the source language per segment. This supports payloads where one input contains multiple source languages; it is not a raw Gemma prompt/tokenizer path.
Returns all known models from merged config together with their live runtime state.
This endpoint is the main UI source of truth.
Example response:
{
"models": [
{
"name": "google_gemma-4-E2B-it-Q8_0-gguf",
"resolved_backend": "llama_cpp",
"configured_enabled": true,
"runtime_state": "loaded",
"is_loaded": true,
"replicas": 3,
"replica_max": 4,
"loaded_replicas": 3,
"inflight_requests": 0,
"queue_depth": 0,
"runtime_inflight": 0,
"configured_target_inflight": 1,
"effective_target_inflight": 1,
"fairness": {
"rejected_per_key_limit": 0,
"rejected_executor_limit": 0,
"keys": []
},
"last_error": null,
"vram_estimate_mib": 57200,
"vram_estimate_replica_count": 3,
"vram_estimate_source": "model_artifact_size",
"load_constraints": {
"gguf_n_ctx": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"gguf_flash_attn": {
"kind": "enum",
"default": "auto",
"allowed_values": ["on", "off", "auto"],
"examples": ["auto", "on", "off"]
},
"gguf_type_k": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
},
"gguf_type_v": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
}
},
"load_recommendations": {
"gguf_cache_type_pairs": {
"kind": "pair_presets",
"fields": ["gguf_type_k", "gguf_type_v"],
"recommended_pairs": [
{
"label": "f16/f16",
"gguf_type_k": "f16",
"gguf_type_v": "f16"
},
{
"label": "q8_0/q8_0",
"gguf_type_k": "q8_0",
"gguf_type_v": "q8_0"
},
{
"label": "q4_0/q4_0",
"gguf_type_k": "q4_0",
"gguf_type_v": "q4_0"
}
],
"notes": [
"Service-curated presets for GGUF cache types.",
"Prefer symmetric GGUF K/V pairs by default; asymmetric pairs may reduce or disable GPU offload in upstream llama.cpp."
]
}
},
"load_override": {},
"capabilities": {
"modalities": ["text"],
"file_inputs": false,
"multi_turn": true,
"thinking_modes": ["default", "enabled", "disabled"],
"response_formats": ["text"]
},
"definition": {
"model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
"backend": "llama_cpp",
"prompt_format": "gemma4_template",
"enable_thinking": null,
"enabled": true,
"replicas": 3,
"replica_max": 4,
"target_inflight": 1,
"gguf_n_gpu_layers": -1,
"gguf_n_ctx": 4096,
"gguf_flash_attn": "auto",
"gguf_type_k": null,
"gguf_type_v": null
}
}
]
}Notes:
configured_enabledreports what the merged config saysruntime_statereports the live process statereplicasreports the current effective replica count for the admin rowdefinition.replicasreports the configured default replica countloaded_replicasreports how many replicas of the public model are currently loadedqueue_depthis the public-model queue depth inside the schedulerruntime_inflightis aggregate inflight work across loaded replicas of the public modelconfigured_target_inflightis the configured per-replica inflight targeteffective_target_inflightis the per-replica scheduler target after capability clamping;llama_server,openai_remote,trtllm_serve,sglang_serve, andvllm_servemay use a configured value above 1, while other backends are clamped to 1fairness.keysreports bounded per-key pending work, active work, configured weight, normalized score, and queue-limit rejection counts; anullkey is the anonymous bucketfairness.rejected_per_key_limitandfairness.rejected_executor_limitare aggregate counters for the current loaded executor and reset on unloadvram_estimate_mibis an approximate per-model VRAM estimatevram_estimate_replica_countis the replica count that the VRAM estimate was measured or derived forvram_estimate_sourceis eitherobserved_load_delta,model_artifact_size, orunavailablecapabilities.modalitieslists the accepted input modalities:text,image, andaudio; a model may advertise any configured combination, withtextadded by defaultcapabilities.file_inputsreports whether the model acceptsfilecontent items; this is currently limited toopenai_remotemodels withremote_file_modeconfiguredcapabilities.multi_turnreports whether the model accepts a multi-turnmessagesarray onPOST /v1/responses; this istrueforllama_server,openai_remote,trtllm_serve,sglang_serve,vllm, andvllm_servemodels and for supported text-onlyllama_cppchat prompt formats (generic,mistral_template,qwen3_template,gemma4_template), but remainsfalseforllama_cpptranslategemma_templatecapabilities.thinking_modeslists accepted values for request-levelthinking; models without a safe per-request control report only["default"], while supported vLLM Gemma4/Qwen3,vllm_serveGemma4, TensorRT-LLM Gemma4, SGLang Gemma4,llama_cppGemma4, ExLlamaV3 Gemma4/Qwen3, CT2 Qwen3, and configured remote models report["default", "enabled", "disabled"]capabilities.reasoning_effortslists provider-defined values accepted byreasoning_effort; an empty list means the model has no such controlcapabilities.thinking_token_budgetisnullwhen unsupported, otherwise it contains the inclusive{ "minimum", "maximum" }range; a request must still leave at least one output token after the budgetcapabilities.response_formatsis["text", "json_schema"]forvllm_serve;trtllm_serve,sglang_serve, and other backends report["text"]load_constraintsdescribes backend-specific live-load fields for UI controlsload_recommendationsdescribes service-curated recommended presets and pairings for UI defaultsload_overridereports the runtime-only override currently active on a loaded modeldefinitioncontains common model fields plus only the fields relevant to the resolved backend
For UI work, load_constraints is the source of truth for which live-load controls should be shown for a model.
Rules:
- if a field is absent from
load_constraints, the UI should treat that field as unsupported for that model - for
kind: "integer", the UI should useminimumandstepdirectly for numeric inputs or sliders - for
kind: "enum", the UI should useallowed_valuesdirectly for a constrained select or segmented control - for
kind: "string_or_null", the UI should use a text input or a constrained select if the frontend chooses to offer known values - if a
defaultis present inload_constraints, the UI may use it as the concrete runtime default when bothdefinitionandload_overrideresolve tonull load_constraintsis derived from the resolved backend, not from whether the model is currently loaded or unloaded- when this document and upstream backend docs differ, the UI should follow the live
load_constraintspayload returned by the service
For UI work, load_recommendations is the source of truth for which presets the service recommends surfacing first.
Rules:
load_recommendationsis optional and additive; it does not replaceload_constraints- fields listed in
recommended_pairsmust still be validated againstload_constraints - the service may accept more combinations than it recommends
- the UI should treat these presets as convenience defaults, not as an exhaustive list of allowed values
For UI state, definition and load_override should be interpreted together:
definitionis the configured value from merged configload_overrideis a runtime-only sparse patch- the effective loaded value is computed by applying
load_overrideoverdefinition - key presence in
load_overridematters, even when the value isnull
This means the UI should not use truthiness to merge values.
Correct merge rule:
if key exists in load_override:
effective_value = load_override[key]
else:
effective_value = definition[key]
This matters in particular for exllama_cache_quant.
Example:
{
"load_override": {
"exllama_cache_quant": null
},
"definition": {
"exllama_cache_quant": "8,8"
}
}In this case, the effective loaded value is null.
For llama_cpp GGUF cache fields, the UI may interpret the effective cache type value as:
- field absent in
load_overrideand absent ornullindefinition: useload_constraints.<field>.default, currently"f16" - effective value
null: useload_constraints.<field>.default, currently"f16" - effective value
"q8_0":q8_0 - effective value
"q4_0":q4_0
For ExLlamaV3, the UI may interpret the effective quant value as:
- field absent in
load_overrideand absent ornullindefinition: fp16 - effective value
null: fp16 - effective value
"8":k=8,v=8 - effective value
"8,4":k=8,v=4
The API does not currently return separate k_bits and v_bits fields.
The UI should parse exllama_cache_quant itself when it wants to display separate K/V values.
Backends that support a load-time inflight target include this constraint:
{
"target_inflight": {
"kind": "integer",
"minimum": 1,
"step": 1
}
}The examples below show the additional backend-specific constraints.
llama_cpp GGUF:
{
"gguf_n_ctx": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"gguf_flash_attn": {
"kind": "enum",
"default": "auto",
"allowed_values": ["on", "off", "auto"],
"examples": ["auto", "on", "off"]
},
"gguf_type_k": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
},
"gguf_type_v": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
}
}GGUF recommended presets:
{
"gguf_cache_type_pairs": {
"kind": "pair_presets",
"fields": ["gguf_type_k", "gguf_type_v"],
"recommended_pairs": [
{
"label": "f16/f16",
"gguf_type_k": "f16",
"gguf_type_v": "f16"
},
{
"label": "q8_0/q8_0",
"gguf_type_k": "q8_0",
"gguf_type_v": "q8_0"
},
{
"label": "q4_0/q4_0",
"gguf_type_k": "q4_0",
"gguf_type_v": "q4_0"
}
]
}
}ExLlamaV3:
{
"exllama_cache_size": {
"kind": "integer",
"minimum": 256,
"step": 256
},
"exllama_max_rq_tokens": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"exllama_cache_k_bits": {
"kind": "integer_or_null",
"minimum": 2,
"maximum": 8,
"default": null,
"null_means": "fp16",
"allowed_values": [2, 3, 4, 5, 6, 7, 8]
},
"exllama_cache_v_bits": {
"kind": "integer_or_null",
"minimum": 2,
"maximum": 8,
"default": null,
"null_means": "fp16",
"allowed_values": [2, 3, 4, 5, 6, 7, 8]
},
"exllama_cache_quant": {
"kind": "string_or_null",
"format": "<bits>|<k_bits>,<v_bits>"
}
}vLLM and vLLM Serve:
{
"vllm_max_model_len": {
"kind": "integer",
"minimum": 256,
"step": 256
},
"vllm_kv_cache_dtype": {
"kind": "enum",
"default": "auto",
"allowed_values": ["auto", "fp8", "fp8_e4m3", "fp8_e5m2"],
"examples": ["auto", "fp8"]
},
"vllm_kv_cache_memory_bytes": {
"kind": "integer",
"minimum": 268435456,
"step": 268435456,
"unit": "bytes",
"display_unit": "mib"
},
"vllm_max_pixels": {
"kind": "integer",
"minimum": 200704,
"step": 200704,
"unit": "pixels"
},
"vllm_speculative_method": {
"kind": "string_or_null",
"format": "vllm_speculative_method",
"default": null,
"examples": ["mtp", "draft_model", "mlp_speculator"]
},
"vllm_speculative_model": {
"kind": "string_or_null",
"format": "hf_id_or_local_path",
"default": null,
"examples": ["google/gemma-4-26B-A4B-it-assistant"]
},
"vllm_speculative_moe_backend": {
"kind": "string_or_null",
"format": "vllm_moe_backend",
"default": null,
"examples": ["triton", "marlin"]
},
"vllm_speculative_attention_backend": {
"kind": "string_or_null",
"format": "vllm_attention_backend",
"default": null,
"examples": ["triton_attn", "flashinfer"]
},
"vllm_num_speculative_tokens": {
"kind": "integer",
"minimum": 1,
"step": 1,
"default": 1
}
}TensorRT-LLM Serve:
{
"trtllm_max_seq_len": {
"kind": "integer",
"minimum": 256,
"step": 256
},
"trtllm_kv_cache_memory_bytes": {
"kind": "integer",
"minimum": 268435456,
"step": 268435456,
"unit": "bytes",
"display_unit": "mib"
},
"trtllm_max_num_tokens": {
"kind": "integer",
"minimum": 256,
"step": 256
},
"trtllm_enable_chunked_prefill": {
"kind": "boolean",
"default": false
},
"trtllm_kv_cache_dtype": {
"kind": "enum",
"default": "auto",
"allowed_values": ["auto", "fp8", "nvfp4"],
"examples": ["auto", "fp8", "nvfp4"]
}
}SGLang Serve:
{
"sglang_context_length": {
"kind": "integer",
"minimum": 256,
"step": 256
},
"sglang_mem_fraction_static": {
"kind": "float",
"minimum": 0.01,
"maximum": 1.0,
"step": 0.01
},
"sglang_max_total_tokens": {
"kind": "integer",
"minimum": 256,
"step": 256,
"unit": "tokens"
},
"sglang_chunked_prefill_size": {
"kind": "integer",
"minimum": -1,
"step": 1
},
"sglang_kv_cache_dtype": {
"kind": "enum",
"default": "auto",
"allowed_values": [
"auto",
"bf16",
"bfloat16",
"fp8_e4m3",
"fp8_e5m2",
"mxfp8",
"nvfp4",
"fp4_mx_block16",
"fp4_e2m1"
]
},
"sglang_speculative_algorithm": {
"kind": "string_or_null",
"format": "sglang_speculative_algorithm",
"default": null,
"examples": ["NEXTN"]
},
"sglang_speculative_draft_model": {
"kind": "string_or_null",
"format": "hf_id_or_local_path",
"default": null,
"examples": ["google/gemma-4-26B-A4B-it-assistant"]
},
"sglang_speculative_num_steps": {
"kind": "integer",
"minimum": 1,
"step": 1,
"default": 5
},
"sglang_speculative_num_draft_tokens": {
"kind": "integer",
"minimum": 1,
"step": 1,
"default": 6
},
"sglang_speculative_eagle_topk": {
"kind": "integer",
"minimum": 1,
"step": 1,
"default": 1
}
}llama-server:
{
"llama_server_n_ctx": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"llama_server_image_max_tokens": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"llama_server_spec_type": {
"kind": "enum",
"default": "draft-mtp",
"allowed_values": ["draft-mtp"],
"examples": ["draft-mtp"]
},
"llama_server_spec_draft_n_max": {
"kind": "integer",
"minimum": 1,
"maximum": 6,
"step": 1,
"default": 2
},
"llama_server_spec_draft_p_min": {
"kind": "float",
"minimum": 0.0,
"maximum": 1.0,
"default": 0.0
}
}ExLlamaV3 recommended presets:
{
"exllama_cache_bit_pairs": {
"kind": "pair_presets",
"fields": ["exllama_cache_k_bits", "exllama_cache_v_bits"],
"recommended_pairs": [
{
"label": "fp16",
"exllama_cache_k_bits": null,
"exllama_cache_v_bits": null
},
{
"label": "8/8",
"exllama_cache_k_bits": 8,
"exllama_cache_v_bits": 8
},
{
"label": "8/4",
"exllama_cache_k_bits": 8,
"exllama_cache_v_bits": 4
}
]
}
}CT2, openai_remote, and stub:
{}Returns current GPU memory usage (from nvidia-smi) and per-model VRAM estimates.
Example response:
{
"gpus": [
{
"index": 0,
"name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition",
"used_mib": 75603,
"total_mib": 97887,
"used_over_total": "75603MiB / 97887MiB"
}
],
"models": [
{
"name": "google_gemma-4-E2B-it-Q8_0-gguf",
"runtime_state": "loaded",
"is_loaded": true,
"configured_target_inflight": 1,
"effective_target_inflight": 1,
"vram_estimate_mib": 12500,
"vram_estimate_replica_count": 3,
"vram_estimate_source": "model_artifact_size"
},
{
"name": "mistral-small-3.2-24b-instruct-2506-gguf",
"runtime_state": "unloaded",
"is_loaded": false,
"configured_target_inflight": 1,
"effective_target_inflight": 1,
"vram_estimate_mib": 16800,
"vram_estimate_replica_count": 1,
"vram_estimate_source": "model_artifact_size"
}
],
"error": null
}Notes:
used_over_totalmatches the compact view you typically read fromnvidia-smivram_estimate_mibfor unloaded models is still an estimate, not a reservationvram_estimate_replica_counttells the caller how many replicas that estimate corresponds to- if
nvidia-smiis unavailable,gpuscan be empty anderrorwill explain why
Rereads the configured settings.json and its matching local.json. The merge rules are the same as at startup.
The operation updates the in-process model catalog. It does not load or unload a model. A newly added model therefore starts in unloaded, even when its file definition has enabled: true.
The reload is rejected with 409 settings_reload_conflict when it would invalidate an active runtime. Conflicts include:
- changing or removing a model that is
loading,loaded, orunloading - changing
engine.decodingorengine.fairnesswhile a model is active
Changing only enabled is safe because that field controls startup loading. The model still appears in updated_models. That list reports file-definition changes, not required runtime reloads. To change another field on a loaded model, unload it first and call the reload endpoint again.
Service changes do not block a catalog reload. The response sets service_restart_required to true until the process restarts with those service settings. This applies to fields such as the bind address, port, and log level.
Invalid JSON or invalid setting values return 400 invalid_settings. A rejected reload leaves the current catalog unchanged.
Example response:
{
"added_models": ["new-model"],
"removed_models": ["retired-model"],
"updated_models": ["changed-model"],
"unchanged_models": ["loaded-model"],
"service_restart_required": false
}Loads one model that already exists in merged config.
Rules:
404ifmodel_nameis unknown200if the model is alreadyloadedorloading- transition
unloaded -> loading -> loaded - transition
failed -> loading -> loaded - if load fails, transition to
failedand retainlast_error - an optional request body may provide
replicasfor this load, but only while the model isunloadedorfailed llama_server,openai_remote,vllm_serve,trtllm_serve, andsglang_servealso accepttarget_inflight- an optional request body may provide temporary backend-specific load overrides for this one live load
Supported load override fields:
- public model:
replicas - supported concurrent server backends:
target_inflight llama_cpp:gguf_n_ctx,gguf_flash_attn,gguf_type_k,gguf_type_v- ExLlamaV3:
exllama_cache_size,exllama_cache_quant,exllama_cache_k_bits,exllama_cache_v_bits,exllama_max_rq_tokens - vLLM and vLLM Serve:
vllm_max_model_len,vllm_kv_cache_dtype,vllm_kv_cache_memory_bytes,vllm_max_pixels,vllm_speculative_method,vllm_speculative_model,vllm_speculative_moe_backend,vllm_speculative_attention_backend,vllm_num_speculative_tokens - TensorRT-LLM Serve:
trtllm_max_seq_len,trtllm_kv_cache_memory_bytes,trtllm_max_num_tokens,trtllm_enable_chunked_prefill,trtllm_kv_cache_dtype - SGLang Serve:
sglang_context_length,sglang_mem_fraction_static,sglang_max_total_tokens,sglang_chunked_prefill_size,sglang_kv_cache_dtype,sglang_speculative_algorithm,sglang_speculative_draft_model,sglang_speculative_num_steps,sglang_speculative_num_draft_tokens,sglang_speculative_eagle_topk - llama-server:
llama_server_n_ctx,llama_server_image_max_tokens,llama_server_spec_type,llama_server_spec_draft_n_max,llama_server_spec_draft_p_min
Example load bodies:
{
"replicas": 3,
"target_inflight": 4
}{
"gguf_n_ctx": 8192
}{
"gguf_n_ctx": 16384
}{
"gguf_n_ctx": 32768
}{
"gguf_n_ctx": 32768,
"gguf_flash_attn": "auto",
"gguf_type_k": "q8_0",
"gguf_type_v": "q4_0"
}{
"exllama_cache_size": 32768,
"exllama_cache_quant": null,
"exllama_max_rq_tokens": 32768
}{
"exllama_cache_size": 32768,
"exllama_cache_quant": "8,8",
"exllama_max_rq_tokens": 32768
}{
"exllama_cache_size": 32768,
"exllama_cache_quant": "8,4",
"exllama_max_rq_tokens": 32768
}{
"exllama_cache_size": 32768,
"exllama_cache_k_bits": 8,
"exllama_cache_v_bits": 4,
"exllama_max_rq_tokens": 32768
}{
"vllm_max_model_len": 16384,
"vllm_kv_cache_dtype": "fp8",
"vllm_kv_cache_memory_bytes": 2147483648,
"vllm_max_pixels": 4014080,
"vllm_speculative_method": "mtp",
"vllm_speculative_model": "google/gemma-4-26B-A4B-it-assistant",
"vllm_speculative_moe_backend": "triton",
"vllm_speculative_attention_backend": "triton_attn",
"vllm_num_speculative_tokens": 1
}{
"target_inflight": 4,
"trtllm_max_seq_len": 20480,
"trtllm_kv_cache_memory_bytes": 8589934592,
"trtllm_max_num_tokens": 8192,
"trtllm_enable_chunked_prefill": false,
"trtllm_kv_cache_dtype": "fp8"
}{
"target_inflight": 4,
"sglang_context_length": 20480,
"sglang_mem_fraction_static": 0.35,
"sglang_max_total_tokens": 20480,
"sglang_chunked_prefill_size": 8192,
"sglang_kv_cache_dtype": "fp8_e4m3",
"sglang_speculative_algorithm": "NEXTN",
"sglang_speculative_draft_model": "google/gemma-4-26B-A4B-it-assistant",
"sglang_speculative_num_steps": 5,
"sglang_speculative_num_draft_tokens": 6,
"sglang_speculative_eagle_topk": 1
}{
"llama_server_n_ctx": 4096,
"llama_server_image_max_tokens": 512,
"llama_server_spec_type": "draft-mtp",
"llama_server_spec_draft_n_max": 4,
"llama_server_spec_draft_p_min": 0.25
}vLLM and vLLM Serve load override notes:
- For
vllm_serve,target_inflightmaps to--max-num-seqsand also controls llm-pool admission. vllm_max_model_lenis the per-load context length.vllm_kv_cache_dtypequantizes the KV cache; allowed UI values areauto,fp8,fp8_e4m3,fp8_e5m2. The service accepts any dtype string vLLM supports.vllm_kv_cache_memory_bytessets an absolute KV cache size in bytes. It is machine-independent and overridesvllm_gpu_memory_utilizationfor KV sizing. Theload_constraintsentry carriesunit: "bytes"anddisplay_unit: "mib"so the UI can present it in MiB.- Prefer keeping configured
vllm_gpu_memory_utilizationvery low and controlling load-time cache budget withvllm_kv_cache_memory_bytes, otherwise vLLM may reserve most free VRAM. vllm_max_pixelscaps the vision-token budget per image for vision-language models; it is merged into the model'svllm_mm_processor_kwargsasmax_pixels.vllm_speculative_methodselects the vLLM speculative path for this load, for examplemtp,draft_model, ormlp_speculator.nulldisables the configured speculative path for that load.vllm_speculative_modelis the assistant/draft/speculator checkpoint or local path passed through vLLM'sspeculative_config.model. For Gemma 4 MTP this is the Gemma 4 assistant checkpoint, not a generic smaller draft model.vllm_speculative_moe_backendmaps to vLLMspeculative_config.moe_backend.nullclears the configured override for this load.vllm_speculative_attention_backendmaps to vLLMspeculative_config.attention_backend.nullclears the configured override for this load.vllm_num_speculative_tokensmaps to vLLMspeculative_config.num_speculative_tokens.- For
vllm_serve, target model id/path, binary path, library path, environment, host, port, API key, and extra CLI args are configured in the model definition, not overridden through the admin load body. vllm_serve_extra_argsmust not set--max-num-seqs;target_inflightowns it.- Loading a
vllm_servemodel starts a localvllm servesubprocess. Unloading terminates that subprocess, so VRAM is released by the server process rather than by Python object cleanup alone.
TensorRT-LLM Serve load notes:
target_inflightmaps to TensorRT-LLMmax_batch_sizeand also controls llm-pool admission.trtllm_max_seq_len,trtllm_max_num_tokens, andtrtllm_enable_chunked_prefillmap to the same top-level TensorRT-LLM YAML fields.trtllm_kv_cache_memory_bytesmaps tokv_cache_config.max_gpu_total_bytes. The API uses bytes;load_constraintstells clients to display MiB.trtllm_kv_cache_dtypemaps tokv_cache_config.dtypeand acceptsauto,fp8, ornvfp4.- llm-pool reads the configured base YAML, merges the effective values and the
target_inflightbatch size into a temporary YAML file, and passes that file through--config. The source YAML is unchanged. Normal runtime cleanup removes the temporary file. - TensorRT-LLM limits KV-cache memory to the lower result of
max_gpu_total_bytesandfree_gpu_memory_fraction. Keep the fractional value high enough to act only as a safety ceiling when an absolute budget should control allocation. This differs from the vLLM pattern of configuring a deliberately low utilization fraction alongside an absolute cache size. - Target model, binary, library path, environment, base YAML, parser names, host, port, timeouts, and extra CLI args remain in the model definition. Neither the base YAML nor
trtllm_serve_extra_argsmay setmax_batch_size. - Loading starts a local
trtllm-serveprocess group. Unloading terminates that group, so VRAM is released by process exit.
SGLang Serve load notes:
target_inflightis passed to SGLang as--max-running-requestsand also controls llm-pool admission. SGLang may reduce its native limit during KV-cache sizing; excess admitted requests then wait inside SGLang.sglang_context_length,sglang_max_total_tokens,sglang_chunked_prefill_size, andsglang_kv_cache_dtypemap to the corresponding SGLang server flags.sglang_max_total_tokenssets the absolute KV-cache token capacity.sglang_mem_fraction_staticremains a startup safety ceiling and must still be high enough to hold the target and assistant weights.sglang_speculative_algorithm: "NEXTN"with a Gemma 4 assistant checkpoint selects SGLang's Frozen-KV MTP path. Set the algorithm tonullto disable speculative decoding for one load.- The speculative step count and top-k map directly to SGLang server flags. When top-k is 1, llm-pool derives the draft-token count as the step count plus 1 and passes it explicitly; a conflicting explicit override is rejected.
sglang_serve_extra_argsmust not set--max-running-requests;target_inflightowns it.- Target model, binary, library path, environment, quantization, attention backend, host, port, parser names, timeouts, and extra CLI arguments remain in the model definition.
- Loading starts a local
sglang serveprocess group. Unloading terminates that group, so VRAM is released by process exit.
llama-server load override notes:
target_inflightmaps to llama-server--paralleland also controls llm-pool admission.llama_server_n_ctxmaps to the nativellama-server -c/--ctx-sizeflag for this load.- llama-server shares
llama_server_n_ctxacross its parallel slots. The approximate per-request context limit isllama_server_n_ctx / target_inflight. llama_server_image_max_tokensmaps to native--image-max-tokensand controls the per-image vision token budget.llama_server_spec_typecurrently accepts only"draft-mtp"ornull.llama_server_spec_draft_n_maxmaps to native--spec-draft-n-max; the API constrains it to1..6.llama_server_spec_draft_p_minmaps to native--spec-draft-p-min; the API constrains it to0.0..1.0.- Model path, binary path, library path,
mmproj, draft model path, GPU layers, flash attention, reasoning, host, port, API key, and extra native args are configured in the model definition, not overridden through the admin load body. llama_server_extra_argsmust not set--parallel;target_inflightowns it.- Loading a
llama_servermodel starts a localllama-serversubprocess. Unloading terminates that subprocess, so VRAM is released by the native server process rather than by Python object cleanup alone.
exllama_cache_quant format:
- omitted or
null: fp16 KV cache "<bits>": same quantization for K and V, for example"8""<k_bits>,<v_bits>": separate K/V quantization, for example"8,4"
exllama_cache_k_bits and exllama_cache_v_bits format:
- both omitted: do not override the current configured value
- both
null: reset to fp16 KV cache - both integers from
2through8: override K and V separately - they must be provided together
- they cannot be combined with
exllama_cache_quantin the same load request
gguf_type_k and gguf_type_v format:
- omitted or
null: use the runtime default cache type "<ggml_type_name>": a GGML cache type name, for example"f16","q8_0", or"q4_0"
gguf_flash_attn format:
- omitted: do not override the current configured value
"on": force Flash Attention on"off": force Flash Attention off"auto": use the runtime auto mode
Upstream references for these backend-specific value sets:
llama_cppGGUF cacheallowed_valuesand defaultf16are based on the officialllama.cppserver docs: https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md- ExLlamaV3
k_bits/v_bitsallowed range2..8is based on the official ExLlamaV3 README, which documents2-8 bit cache quantization: https://github.com/turboderp-org/exllamav3 - The conservative GGUF preset guidance above is informed by upstream llama.cpp reports that asymmetric K/V pairs can disable GPU offload in some setups: ggml-org/llama.cpp#20866
These overrides are runtime-only:
- they do not modify
settings.json - they do not modify
local.json - they are returned separately from the configured definition in admin responses
replicasin the load request does not modifydefinition.replicas; it only selects the replica count for that one live loadtarget_inflightin the load request does not modifydefinition.target_inflight; it sets scheduler and managed-server concurrency for that one live load
Example response:
{
"name": "google_gemma-4-E2B-it-Q8_0-gguf",
"resolved_backend": "llama_cpp",
"configured_enabled": false,
"runtime_state": "loaded",
"is_loaded": true,
"replicas": 3,
"replica_max": 4,
"loaded_replicas": 3,
"inflight_requests": 0,
"queue_depth": 0,
"runtime_inflight": 0,
"configured_target_inflight": 1,
"effective_target_inflight": 1,
"fairness": {
"rejected_per_key_limit": 0,
"rejected_executor_limit": 0,
"keys": []
},
"last_error": null,
"vram_estimate_mib": 12340,
"vram_estimate_replica_count": 3,
"vram_estimate_source": "observed_load_delta",
"load_constraints": {
"gguf_n_ctx": {
"kind": "integer",
"minimum": 1,
"step": 1
},
"gguf_flash_attn": {
"kind": "enum",
"default": "auto",
"allowed_values": ["on", "off", "auto"],
"examples": ["auto", "on", "off"]
},
"gguf_type_k": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
},
"gguf_type_v": {
"kind": "string_or_null",
"format": "ggml_type_name",
"default": "f16",
"allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
"examples": ["f16", "q8_0", "q4_0"]
}
},
"load_recommendations": {
"gguf_cache_type_pairs": {
"kind": "pair_presets",
"fields": ["gguf_type_k", "gguf_type_v"],
"recommended_pairs": [
{
"label": "f16/f16",
"gguf_type_k": "f16",
"gguf_type_v": "f16"
},
{
"label": "q8_0/q8_0",
"gguf_type_k": "q8_0",
"gguf_type_v": "q8_0"
},
{
"label": "q4_0/q4_0",
"gguf_type_k": "q4_0",
"gguf_type_v": "q4_0"
}
],
"notes": [
"Service-curated presets for GGUF cache types.",
"Prefer symmetric GGUF K/V pairs by default; asymmetric pairs may reduce or disable GPU offload in upstream llama.cpp."
]
}
},
"load_override": {
"gguf_n_ctx": 32768,
"gguf_flash_attn": "auto",
"gguf_type_k": "q8_0",
"gguf_type_v": "q4_0"
},
"capabilities": {
"modalities": ["text"],
"file_inputs": false,
"multi_turn": true,
"thinking_modes": ["default", "enabled", "disabled"],
"response_formats": ["text"]
},
"definition": {
"model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
"backend": "llama_cpp",
"prompt_format": "gemma4_template",
"enable_thinking": null,
"enabled": false,
"replicas": 3,
"replica_max": 4,
"target_inflight": 1,
"gguf_n_gpu_layers": -1,
"gguf_n_ctx": 4096,
"gguf_flash_attn": "auto",
"gguf_type_k": null,
"gguf_type_v": null
}
}Notes:
- loading is allowed for configured models even when
configured_enabledisfalse - after a successful load,
vram_estimate_sourcemay switch toobserved_load_deltaif a GPU delta could be measured during load - when
replicasis provided, the model load is aggregate and all-or-nothing for that selected replica count replicasand load overrides may only be changed while the model isunloadedorfailed
Validation behavior:
422means the request body failed schema validation before runtime logic ran Examples:target_inflight: 0gguf_n_ctx: 0exllama_cache_size: 0exllama_max_rq_tokens: 0trtllm_kv_cache_memory_bytes: 0400withcode: "invalid_load_request"means the body was structurally valid, but the values were invalid for the resolved backend or runtime rules Examples:gguf_type_k: "q8-0"gguf_type_k: "foo"sending onlyexllama_cache_k_bitswithoutexllama_cache_v_bitscombiningexllama_cache_quantwithexllama_cache_k_bits/exllama_cache_v_bitsexllama_cache_size: 8000exllama_cache_quant: "fp16"llama_server_spec_type: "medusa"llama_server_spec_draft_n_max: 7llama_server_spec_draft_p_min: 1.5trtllm_kv_cache_dtype: "int8"sending ExLlamaV3-only fields to a llama_cpp model sending llama-server-only fields to a vLLM model sending load overrides while the model is already loaded and not first unloading it409still applies for runtime state conflicts such as loading or unloading transitions Examples: loading a model while it is already unloading
Gracefully unloads one currently loaded model.
Rules:
404ifmodel_nameis unknown200if alreadyunloaded200if alreadyfailed200if alreadyunloading- transition
loaded -> unloading -> unloaded - unload stops all loaded replicas of the public model
- new inference requests are rejected once
unloadingstarts - in-flight requests are allowed to finish before resources are released
Selected fields from an unloaded response; the complete object has the same
AdminModelEntry shape as an item returned by GET /v1/admin/models:
{
"name": "google_gemma-4-E2B-it-Q8_0-gguf",
"resolved_backend": "llama_cpp",
"configured_enabled": false,
"runtime_state": "unloaded",
"is_loaded": false,
"replicas": 3,
"replica_max": 4,
"loaded_replicas": 0,
"inflight_requests": 0,
"queue_depth": 0,
"runtime_inflight": 0,
"configured_target_inflight": 1,
"effective_target_inflight": 1,
"fairness": {
"rejected_per_key_limit": 0,
"rejected_executor_limit": 0,
"keys": []
},
"last_error": null,
"vram_estimate_mib": 12340,
"vram_estimate_replica_count": 4,
"vram_estimate_source": "observed_load_delta",
"definition": {
"model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
"backend": "llama_cpp",
"device": "cuda",
"prompt_format": "gemma4_template",
"enabled": false,
"replicas": 3,
"replica_max": 4,
"target_inflight": 1
}
}Unload is graceful.
The service:
- marks the model as
unloading - rejects new inference requests for that model
- cancels requests still waiting in the scheduler queue
- waits for active backend requests to finish
- releases runtime references and backend resources
- marks the model as
unloaded
inflight_requests tracks admitted requests so unload can wait safely while
queued requests are cancelled and active requests drain.
The scheduler owns pending request queues. The runtime owns backend execution state.
On unload, the scheduler and runtime therefore:
- stop new admissions
- cancel queued work with
model_unloading - let active backend work drain
- unload the model after no admitted requests remain
A successful unload means:
- no runtime remains registered for the model
- no new requests can reach that runtime
- no in-flight requests remain
- backend-owned objects are dereferenced
- memory becomes reusable for later loads
The API does not promise that every allocator reports zero immediately after unload.
The guarantee is functional reuse, not cosmetic memory counters.
The API supports UI clients through:
- stable response models
- OpenAPI descriptions on every admin endpoint
- clear descriptions of runtime states
- clear error codes for rejected inference requests
The UI should be able to render:
- configured vs loaded state
- current lifecycle state
- last load error
The admin API does not define:
- writes back to config files
- force unload
- active GPU job cancellation
- retry policy for failed model loads