Skip to content

Latest commit

 

History

History
1329 lines (1138 loc) · 45.7 KB

File metadata and controls

1329 lines (1138 loc) · 45.7 KB

Runtime Admin API

This document describes the implemented admin API for loading, unloading, and inspecting models at runtime.

The API avoids edits to local.json and service restarts for routine model management.

It controls live runtime state only:

  • no automatic writes back to settings.json or local.json
  • no arbitrary model definitions via API
  • no force unload
  • no background job system for model loads

Managed backend behavior:

  • this admin API is implemented and is the live control plane used by the workbench
  • replicas is a common runtime-only load override; target_inflight is available for llama_server, openai_remote, vllm_serve, trtllm_serve, and sglang_serve; backend-specific overrides are implemented for llama_cpp, exllamav3, vllm, vllm_serve, trtllm_serve, sglang_serve, and llama_server
  • llama_server load/unload starts and stops a managed native llama-server subprocess; binary path, model path, library path, mmproj, draft model path, host, port, and extra native args stay in model config
  • vllm_serve load/unload starts and stops a managed local vllm serve subprocess; binary path, target model id/path, library path, environment, host, port, API key, and extra CLI args stay in model config
  • trtllm_serve load/unload starts and stops a managed local trtllm-serve process group; binary path, target model id/path, library path, environment, host, port, TensorRT-LLM config, parser names, and extra CLI args stay in model config
  • sglang_serve load/unload starts and stops a managed local sglang serve process group; binary path, target model id/path, library path, environment, host, port, parser names, and extra CLI args stay in model config
  • clients use the live load_constraints payload to build backend-specific load controls

Contents

Purpose

The current service merges settings.json and local.json into one effective config, then loads enabled models at startup.

The admin API adds a separate live control plane on top of that merged config:

  • the merged config tells us which models are known to the service
  • the live runtime state tells us which of those models are currently loaded

That distinction must stay explicit in both the API and the UI.

Core Concepts

Configured Model Definition

A configured model definition comes from the merged settings.json + local.json payload.

The process reads this definition at startup and when the settings reload endpoint is called. It includes fields such as:

  • model_path
  • backend
  • device
  • prompt_format
  • backend-specific settings
  • enabled

The admin API does not write this definition. The reload endpoint only rereads the files.

Runtime State

Each configured model also has a live runtime state inside the process.

Allowed states:

  • unloaded
  • loading
  • loaded
  • unloading
  • failed

These states are runtime-only and may differ from the original enabled value in config.

State Semantics

unloaded

  • the model exists in merged config
  • no runtime is currently loaded
  • inference requests for this model are rejected
  • the model may be loaded through the admin API

loading

  • a runtime load has started but is not complete yet
  • inference requests for this model are rejected
  • duplicate load requests without load overrides are idempotent and return the current state

loaded

  • a runtime exists and may serve inference requests
  • the model may be unloaded through the admin API

unloading

  • no new inference requests are accepted for this model
  • in-flight requests are allowed to finish
  • once in-flight requests reach zero, runtime resources are released

failed

  • the last load attempt failed
  • last_error is retained for inspection
  • the model may be loaded again through the admin API

Request Behavior By Runtime State

For POST /v1/responses:

  • loaded: accept
  • unloaded: reject
  • loading: reject
  • unloading: reject
  • failed: reject

Errors are explicit and machine-readable.

Request error codes include:

  • unknown_model
  • model_not_loaded
  • model_loading
  • model_unloading
  • model_failed
  • remote_execution_disallowed
  • file_input_unsupported
  • modality_unsupported
  • thinking_unsupported
  • response_format_unsupported
  • fairness_key_queue_full
  • executor_queue_full

Requests may optionally include thinking: "default" | "enabled" | "disabled". default preserves the model configuration. enabled and disabled are accepted only when the selected model advertises those values in capabilities.thinking_modes; otherwise the request is rejected with 400 thinking_unsupported.

response_format accepts strict JSON Schema output for non-streaming vllm_serve requests. Other backends reject it with 400 response_format_unsupported. Models advertise accepted formats through capabilities.response_formats.

TranslateGemma Request Notes

llama_cpp models configured with prompt_format: "translategemma_template" use the official structured TranslateGemma request shape internally. They remain single-turn text models and continue to report capabilities.multi_turn: false.

Known-source requests should include both source_lang_code and target_lang_code.

Mixed-source requests may omit source_lang_code, or set it to "auto" or "mixed", while still providing target_lang_code. In that mode, the runtime keeps the structured TranslateGemma path, uses an internal valid source-language fallback, and prepends a short instruction asking the model to detect the source language per segment. This supports payloads where one input contains multiple source languages; it is not a raw Gemma prompt/tokenizer path.

Endpoints

GET /v1/admin/models

Returns all known models from merged config together with their live runtime state.

This endpoint is the main UI source of truth.

Example response:

{
  "models": [
    {
      "name": "google_gemma-4-E2B-it-Q8_0-gguf",
      "resolved_backend": "llama_cpp",
      "configured_enabled": true,
      "runtime_state": "loaded",
      "is_loaded": true,
      "replicas": 3,
      "replica_max": 4,
      "loaded_replicas": 3,
      "inflight_requests": 0,
      "queue_depth": 0,
      "runtime_inflight": 0,
      "configured_target_inflight": 1,
      "effective_target_inflight": 1,
      "fairness": {
        "rejected_per_key_limit": 0,
        "rejected_executor_limit": 0,
        "keys": []
      },
      "last_error": null,
      "vram_estimate_mib": 57200,
      "vram_estimate_replica_count": 3,
      "vram_estimate_source": "model_artifact_size",
      "load_constraints": {
        "gguf_n_ctx": {
          "kind": "integer",
          "minimum": 1,
          "step": 1
        },
        "gguf_flash_attn": {
          "kind": "enum",
          "default": "auto",
          "allowed_values": ["on", "off", "auto"],
          "examples": ["auto", "on", "off"]
        },
        "gguf_type_k": {
          "kind": "string_or_null",
          "format": "ggml_type_name",
          "default": "f16",
          "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
          "examples": ["f16", "q8_0", "q4_0"]
        },
        "gguf_type_v": {
          "kind": "string_or_null",
          "format": "ggml_type_name",
          "default": "f16",
          "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
          "examples": ["f16", "q8_0", "q4_0"]
        }
      },
      "load_recommendations": {
        "gguf_cache_type_pairs": {
          "kind": "pair_presets",
          "fields": ["gguf_type_k", "gguf_type_v"],
          "recommended_pairs": [
            {
              "label": "f16/f16",
              "gguf_type_k": "f16",
              "gguf_type_v": "f16"
            },
            {
              "label": "q8_0/q8_0",
              "gguf_type_k": "q8_0",
              "gguf_type_v": "q8_0"
            },
            {
              "label": "q4_0/q4_0",
              "gguf_type_k": "q4_0",
              "gguf_type_v": "q4_0"
            }
          ],
          "notes": [
            "Service-curated presets for GGUF cache types.",
            "Prefer symmetric GGUF K/V pairs by default; asymmetric pairs may reduce or disable GPU offload in upstream llama.cpp."
          ]
        }
      },
      "load_override": {},
      "capabilities": {
        "modalities": ["text"],
        "file_inputs": false,
        "multi_turn": true,
        "thinking_modes": ["default", "enabled", "disabled"],
        "response_formats": ["text"]
      },
      "definition": {
        "model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
        "backend": "llama_cpp",
        "prompt_format": "gemma4_template",
        "enable_thinking": null,
        "enabled": true,
        "replicas": 3,
        "replica_max": 4,
        "target_inflight": 1,
        "gguf_n_gpu_layers": -1,
        "gguf_n_ctx": 4096,
        "gguf_flash_attn": "auto",
        "gguf_type_k": null,
        "gguf_type_v": null
      }
    }
  ]
}

Notes:

  • configured_enabled reports what the merged config says
  • runtime_state reports the live process state
  • replicas reports the current effective replica count for the admin row
  • definition.replicas reports the configured default replica count
  • loaded_replicas reports how many replicas of the public model are currently loaded
  • queue_depth is the public-model queue depth inside the scheduler
  • runtime_inflight is aggregate inflight work across loaded replicas of the public model
  • configured_target_inflight is the configured per-replica inflight target
  • effective_target_inflight is the per-replica scheduler target after capability clamping; llama_server, openai_remote, trtllm_serve, sglang_serve, and vllm_serve may use a configured value above 1, while other backends are clamped to 1
  • fairness.keys reports bounded per-key pending work, active work, configured weight, normalized score, and queue-limit rejection counts; a null key is the anonymous bucket
  • fairness.rejected_per_key_limit and fairness.rejected_executor_limit are aggregate counters for the current loaded executor and reset on unload
  • vram_estimate_mib is an approximate per-model VRAM estimate
  • vram_estimate_replica_count is the replica count that the VRAM estimate was measured or derived for
  • vram_estimate_source is either observed_load_delta, model_artifact_size, or unavailable
  • capabilities.modalities lists the accepted input modalities: text, image, and audio; a model may advertise any configured combination, with text added by default
  • capabilities.file_inputs reports whether the model accepts file content items; this is currently limited to openai_remote models with remote_file_mode configured
  • capabilities.multi_turn reports whether the model accepts a multi-turn messages array on POST /v1/responses; this is true for llama_server, openai_remote, trtllm_serve, sglang_serve, vllm, and vllm_serve models and for supported text-only llama_cpp chat prompt formats (generic, mistral_template, qwen3_template, gemma4_template), but remains false for llama_cpp translategemma_template
  • capabilities.thinking_modes lists accepted values for request-level thinking; models without a safe per-request control report only ["default"], while supported vLLM Gemma4/Qwen3, vllm_serve Gemma4, TensorRT-LLM Gemma4, SGLang Gemma4, llama_cpp Gemma4, ExLlamaV3 Gemma4/Qwen3, CT2 Qwen3, and configured remote models report ["default", "enabled", "disabled"]
  • capabilities.reasoning_efforts lists provider-defined values accepted by reasoning_effort; an empty list means the model has no such control
  • capabilities.thinking_token_budget is null when unsupported, otherwise it contains the inclusive { "minimum", "maximum" } range; a request must still leave at least one output token after the budget
  • capabilities.response_formats is ["text", "json_schema"] for vllm_serve; trtllm_serve, sglang_serve, and other backends report ["text"]
  • load_constraints describes backend-specific live-load fields for UI controls
  • load_recommendations describes service-curated recommended presets and pairings for UI defaults
  • load_override reports the runtime-only override currently active on a loaded model
  • definition contains common model fields plus only the fields relevant to the resolved backend

UI-Facing load_constraints

For UI work, load_constraints is the source of truth for which live-load controls should be shown for a model.

Rules:

  • if a field is absent from load_constraints, the UI should treat that field as unsupported for that model
  • for kind: "integer", the UI should use minimum and step directly for numeric inputs or sliders
  • for kind: "enum", the UI should use allowed_values directly for a constrained select or segmented control
  • for kind: "string_or_null", the UI should use a text input or a constrained select if the frontend chooses to offer known values
  • if a default is present in load_constraints, the UI may use it as the concrete runtime default when both definition and load_override resolve to null
  • load_constraints is derived from the resolved backend, not from whether the model is currently loaded or unloaded
  • when this document and upstream backend docs differ, the UI should follow the live load_constraints payload returned by the service

UI-Facing load_recommendations

For UI work, load_recommendations is the source of truth for which presets the service recommends surfacing first.

Rules:

  • load_recommendations is optional and additive; it does not replace load_constraints
  • fields listed in recommended_pairs must still be validated against load_constraints
  • the service may accept more combinations than it recommends
  • the UI should treat these presets as convenience defaults, not as an exhaustive list of allowed values

Effective Loaded Values

For UI state, definition and load_override should be interpreted together:

  • definition is the configured value from merged config
  • load_override is a runtime-only sparse patch
  • the effective loaded value is computed by applying load_override over definition
  • key presence in load_override matters, even when the value is null

This means the UI should not use truthiness to merge values.

Correct merge rule:

if key exists in load_override:
  effective_value = load_override[key]
else:
  effective_value = definition[key]

This matters in particular for exllama_cache_quant.

Example:

{
  "load_override": {
    "exllama_cache_quant": null
  },
  "definition": {
    "exllama_cache_quant": "8,8"
  }
}

In this case, the effective loaded value is null.

For llama_cpp GGUF cache fields, the UI may interpret the effective cache type value as:

  • field absent in load_override and absent or null in definition: use load_constraints.<field>.default, currently "f16"
  • effective value null: use load_constraints.<field>.default, currently "f16"
  • effective value "q8_0": q8_0
  • effective value "q4_0": q4_0

For ExLlamaV3, the UI may interpret the effective quant value as:

  • field absent in load_override and absent or null in definition: fp16
  • effective value null: fp16
  • effective value "8": k=8, v=8
  • effective value "8,4": k=8, v=4

The API does not currently return separate k_bits and v_bits fields. The UI should parse exllama_cache_quant itself when it wants to display separate K/V values.

Current load_constraints Shapes

Backends that support a load-time inflight target include this constraint:

{
  "target_inflight": {
    "kind": "integer",
    "minimum": 1,
    "step": 1
  }
}

The examples below show the additional backend-specific constraints.

llama_cpp GGUF:

{
  "gguf_n_ctx": {
    "kind": "integer",
    "minimum": 1,
    "step": 1
  },
  "gguf_flash_attn": {
    "kind": "enum",
    "default": "auto",
    "allowed_values": ["on", "off", "auto"],
    "examples": ["auto", "on", "off"]
  },
  "gguf_type_k": {
    "kind": "string_or_null",
    "format": "ggml_type_name",
    "default": "f16",
    "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
    "examples": ["f16", "q8_0", "q4_0"]
  },
  "gguf_type_v": {
    "kind": "string_or_null",
    "format": "ggml_type_name",
    "default": "f16",
    "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
    "examples": ["f16", "q8_0", "q4_0"]
  }
}

GGUF recommended presets:

{
  "gguf_cache_type_pairs": {
    "kind": "pair_presets",
    "fields": ["gguf_type_k", "gguf_type_v"],
    "recommended_pairs": [
      {
        "label": "f16/f16",
        "gguf_type_k": "f16",
        "gguf_type_v": "f16"
      },
      {
        "label": "q8_0/q8_0",
        "gguf_type_k": "q8_0",
        "gguf_type_v": "q8_0"
      },
      {
        "label": "q4_0/q4_0",
        "gguf_type_k": "q4_0",
        "gguf_type_v": "q4_0"
      }
    ]
  }
}

ExLlamaV3:

{
  "exllama_cache_size": {
    "kind": "integer",
    "minimum": 256,
    "step": 256
  },
  "exllama_max_rq_tokens": {
    "kind": "integer",
    "minimum": 1,
    "step": 1
  },
  "exllama_cache_k_bits": {
    "kind": "integer_or_null",
    "minimum": 2,
    "maximum": 8,
    "default": null,
    "null_means": "fp16",
    "allowed_values": [2, 3, 4, 5, 6, 7, 8]
  },
  "exllama_cache_v_bits": {
    "kind": "integer_or_null",
    "minimum": 2,
    "maximum": 8,
    "default": null,
    "null_means": "fp16",
    "allowed_values": [2, 3, 4, 5, 6, 7, 8]
  },
  "exllama_cache_quant": {
    "kind": "string_or_null",
    "format": "<bits>|<k_bits>,<v_bits>"
  }
}

vLLM and vLLM Serve:

{
  "vllm_max_model_len": {
    "kind": "integer",
    "minimum": 256,
    "step": 256
  },
  "vllm_kv_cache_dtype": {
    "kind": "enum",
    "default": "auto",
    "allowed_values": ["auto", "fp8", "fp8_e4m3", "fp8_e5m2"],
    "examples": ["auto", "fp8"]
  },
  "vllm_kv_cache_memory_bytes": {
    "kind": "integer",
    "minimum": 268435456,
    "step": 268435456,
    "unit": "bytes",
    "display_unit": "mib"
  },
  "vllm_max_pixels": {
    "kind": "integer",
    "minimum": 200704,
    "step": 200704,
    "unit": "pixels"
  },
  "vllm_speculative_method": {
    "kind": "string_or_null",
    "format": "vllm_speculative_method",
    "default": null,
    "examples": ["mtp", "draft_model", "mlp_speculator"]
  },
  "vllm_speculative_model": {
    "kind": "string_or_null",
    "format": "hf_id_or_local_path",
    "default": null,
    "examples": ["google/gemma-4-26B-A4B-it-assistant"]
  },
  "vllm_speculative_moe_backend": {
    "kind": "string_or_null",
    "format": "vllm_moe_backend",
    "default": null,
    "examples": ["triton", "marlin"]
  },
  "vllm_speculative_attention_backend": {
    "kind": "string_or_null",
    "format": "vllm_attention_backend",
    "default": null,
    "examples": ["triton_attn", "flashinfer"]
  },
  "vllm_num_speculative_tokens": {
    "kind": "integer",
    "minimum": 1,
    "step": 1,
    "default": 1
  }
}

TensorRT-LLM Serve:

{
  "trtllm_max_seq_len": {
    "kind": "integer",
    "minimum": 256,
    "step": 256
  },
  "trtllm_kv_cache_memory_bytes": {
    "kind": "integer",
    "minimum": 268435456,
    "step": 268435456,
    "unit": "bytes",
    "display_unit": "mib"
  },
  "trtllm_max_num_tokens": {
    "kind": "integer",
    "minimum": 256,
    "step": 256
  },
  "trtllm_enable_chunked_prefill": {
    "kind": "boolean",
    "default": false
  },
  "trtllm_kv_cache_dtype": {
    "kind": "enum",
    "default": "auto",
    "allowed_values": ["auto", "fp8", "nvfp4"],
    "examples": ["auto", "fp8", "nvfp4"]
  }
}

SGLang Serve:

{
  "sglang_context_length": {
    "kind": "integer",
    "minimum": 256,
    "step": 256
  },
  "sglang_mem_fraction_static": {
    "kind": "float",
    "minimum": 0.01,
    "maximum": 1.0,
    "step": 0.01
  },
  "sglang_max_total_tokens": {
    "kind": "integer",
    "minimum": 256,
    "step": 256,
    "unit": "tokens"
  },
  "sglang_chunked_prefill_size": {
    "kind": "integer",
    "minimum": -1,
    "step": 1
  },
  "sglang_kv_cache_dtype": {
    "kind": "enum",
    "default": "auto",
    "allowed_values": [
      "auto",
      "bf16",
      "bfloat16",
      "fp8_e4m3",
      "fp8_e5m2",
      "mxfp8",
      "nvfp4",
      "fp4_mx_block16",
      "fp4_e2m1"
    ]
  },
  "sglang_speculative_algorithm": {
    "kind": "string_or_null",
    "format": "sglang_speculative_algorithm",
    "default": null,
    "examples": ["NEXTN"]
  },
  "sglang_speculative_draft_model": {
    "kind": "string_or_null",
    "format": "hf_id_or_local_path",
    "default": null,
    "examples": ["google/gemma-4-26B-A4B-it-assistant"]
  },
  "sglang_speculative_num_steps": {
    "kind": "integer",
    "minimum": 1,
    "step": 1,
    "default": 5
  },
  "sglang_speculative_num_draft_tokens": {
    "kind": "integer",
    "minimum": 1,
    "step": 1,
    "default": 6
  },
  "sglang_speculative_eagle_topk": {
    "kind": "integer",
    "minimum": 1,
    "step": 1,
    "default": 1
  }
}

llama-server:

{
  "llama_server_n_ctx": {
    "kind": "integer",
    "minimum": 1,
    "step": 1
  },
  "llama_server_image_max_tokens": {
    "kind": "integer",
    "minimum": 1,
    "step": 1
  },
  "llama_server_spec_type": {
    "kind": "enum",
    "default": "draft-mtp",
    "allowed_values": ["draft-mtp"],
    "examples": ["draft-mtp"]
  },
  "llama_server_spec_draft_n_max": {
    "kind": "integer",
    "minimum": 1,
    "maximum": 6,
    "step": 1,
    "default": 2
  },
  "llama_server_spec_draft_p_min": {
    "kind": "float",
    "minimum": 0.0,
    "maximum": 1.0,
    "default": 0.0
  }
}

ExLlamaV3 recommended presets:

{
  "exllama_cache_bit_pairs": {
    "kind": "pair_presets",
    "fields": ["exllama_cache_k_bits", "exllama_cache_v_bits"],
    "recommended_pairs": [
      {
        "label": "fp16",
        "exllama_cache_k_bits": null,
        "exllama_cache_v_bits": null
      },
      {
        "label": "8/8",
        "exllama_cache_k_bits": 8,
        "exllama_cache_v_bits": 8
      },
      {
        "label": "8/4",
        "exllama_cache_k_bits": 8,
        "exllama_cache_v_bits": 4
      }
    ]
  }
}

CT2, openai_remote, and stub:

{}

GET /v1/admin/gpu-memory

Returns current GPU memory usage (from nvidia-smi) and per-model VRAM estimates.

Example response:

{
  "gpus": [
    {
      "index": 0,
      "name": "NVIDIA RTX PRO 6000 Blackwell Workstation Edition",
      "used_mib": 75603,
      "total_mib": 97887,
      "used_over_total": "75603MiB / 97887MiB"
    }
  ],
  "models": [
    {
      "name": "google_gemma-4-E2B-it-Q8_0-gguf",
      "runtime_state": "loaded",
      "is_loaded": true,
      "configured_target_inflight": 1,
      "effective_target_inflight": 1,
      "vram_estimate_mib": 12500,
      "vram_estimate_replica_count": 3,
      "vram_estimate_source": "model_artifact_size"
    },
    {
      "name": "mistral-small-3.2-24b-instruct-2506-gguf",
      "runtime_state": "unloaded",
      "is_loaded": false,
      "configured_target_inflight": 1,
      "effective_target_inflight": 1,
      "vram_estimate_mib": 16800,
      "vram_estimate_replica_count": 1,
      "vram_estimate_source": "model_artifact_size"
    }
  ],
  "error": null
}

Notes:

  • used_over_total matches the compact view you typically read from nvidia-smi
  • vram_estimate_mib for unloaded models is still an estimate, not a reservation
  • vram_estimate_replica_count tells the caller how many replicas that estimate corresponds to
  • if nvidia-smi is unavailable, gpus can be empty and error will explain why

POST /v1/admin/settings/reload

Rereads the configured settings.json and its matching local.json. The merge rules are the same as at startup.

The operation updates the in-process model catalog. It does not load or unload a model. A newly added model therefore starts in unloaded, even when its file definition has enabled: true.

The reload is rejected with 409 settings_reload_conflict when it would invalidate an active runtime. Conflicts include:

  • changing or removing a model that is loading, loaded, or unloading
  • changing engine.decoding or engine.fairness while a model is active

Changing only enabled is safe because that field controls startup loading. The model still appears in updated_models. That list reports file-definition changes, not required runtime reloads. To change another field on a loaded model, unload it first and call the reload endpoint again.

Service changes do not block a catalog reload. The response sets service_restart_required to true until the process restarts with those service settings. This applies to fields such as the bind address, port, and log level.

Invalid JSON or invalid setting values return 400 invalid_settings. A rejected reload leaves the current catalog unchanged.

Example response:

{
  "added_models": ["new-model"],
  "removed_models": ["retired-model"],
  "updated_models": ["changed-model"],
  "unchanged_models": ["loaded-model"],
  "service_restart_required": false
}

POST /v1/admin/models/{model_name}/load

Loads one model that already exists in merged config.

Rules:

  • 404 if model_name is unknown
  • 200 if the model is already loaded or loading
  • transition unloaded -> loading -> loaded
  • transition failed -> loading -> loaded
  • if load fails, transition to failed and retain last_error
  • an optional request body may provide replicas for this load, but only while the model is unloaded or failed
  • llama_server, openai_remote, vllm_serve, trtllm_serve, and sglang_serve also accept target_inflight
  • an optional request body may provide temporary backend-specific load overrides for this one live load

Supported load override fields:

  • public model: replicas
  • supported concurrent server backends: target_inflight
  • llama_cpp: gguf_n_ctx, gguf_flash_attn, gguf_type_k, gguf_type_v
  • ExLlamaV3: exllama_cache_size, exllama_cache_quant, exllama_cache_k_bits, exllama_cache_v_bits, exllama_max_rq_tokens
  • vLLM and vLLM Serve: vllm_max_model_len, vllm_kv_cache_dtype, vllm_kv_cache_memory_bytes, vllm_max_pixels, vllm_speculative_method, vllm_speculative_model, vllm_speculative_moe_backend, vllm_speculative_attention_backend, vllm_num_speculative_tokens
  • TensorRT-LLM Serve: trtllm_max_seq_len, trtllm_kv_cache_memory_bytes, trtllm_max_num_tokens, trtllm_enable_chunked_prefill, trtllm_kv_cache_dtype
  • SGLang Serve: sglang_context_length, sglang_mem_fraction_static, sglang_max_total_tokens, sglang_chunked_prefill_size, sglang_kv_cache_dtype, sglang_speculative_algorithm, sglang_speculative_draft_model, sglang_speculative_num_steps, sglang_speculative_num_draft_tokens, sglang_speculative_eagle_topk
  • llama-server: llama_server_n_ctx, llama_server_image_max_tokens, llama_server_spec_type, llama_server_spec_draft_n_max, llama_server_spec_draft_p_min

Example load bodies:

{
  "replicas": 3,
  "target_inflight": 4
}
{
  "gguf_n_ctx": 8192
}
{
  "gguf_n_ctx": 16384
}
{
  "gguf_n_ctx": 32768
}
{
  "gguf_n_ctx": 32768,
  "gguf_flash_attn": "auto",
  "gguf_type_k": "q8_0",
  "gguf_type_v": "q4_0"
}
{
  "exllama_cache_size": 32768,
  "exllama_cache_quant": null,
  "exllama_max_rq_tokens": 32768
}
{
  "exllama_cache_size": 32768,
  "exllama_cache_quant": "8,8",
  "exllama_max_rq_tokens": 32768
}
{
  "exllama_cache_size": 32768,
  "exllama_cache_quant": "8,4",
  "exllama_max_rq_tokens": 32768
}
{
  "exllama_cache_size": 32768,
  "exllama_cache_k_bits": 8,
  "exllama_cache_v_bits": 4,
  "exllama_max_rq_tokens": 32768
}
{
  "vllm_max_model_len": 16384,
  "vllm_kv_cache_dtype": "fp8",
  "vllm_kv_cache_memory_bytes": 2147483648,
  "vllm_max_pixels": 4014080,
  "vllm_speculative_method": "mtp",
  "vllm_speculative_model": "google/gemma-4-26B-A4B-it-assistant",
  "vllm_speculative_moe_backend": "triton",
  "vllm_speculative_attention_backend": "triton_attn",
  "vllm_num_speculative_tokens": 1
}
{
  "target_inflight": 4,
  "trtllm_max_seq_len": 20480,
  "trtllm_kv_cache_memory_bytes": 8589934592,
  "trtllm_max_num_tokens": 8192,
  "trtllm_enable_chunked_prefill": false,
  "trtllm_kv_cache_dtype": "fp8"
}
{
  "target_inflight": 4,
  "sglang_context_length": 20480,
  "sglang_mem_fraction_static": 0.35,
  "sglang_max_total_tokens": 20480,
  "sglang_chunked_prefill_size": 8192,
  "sglang_kv_cache_dtype": "fp8_e4m3",
  "sglang_speculative_algorithm": "NEXTN",
  "sglang_speculative_draft_model": "google/gemma-4-26B-A4B-it-assistant",
  "sglang_speculative_num_steps": 5,
  "sglang_speculative_num_draft_tokens": 6,
  "sglang_speculative_eagle_topk": 1
}
{
  "llama_server_n_ctx": 4096,
  "llama_server_image_max_tokens": 512,
  "llama_server_spec_type": "draft-mtp",
  "llama_server_spec_draft_n_max": 4,
  "llama_server_spec_draft_p_min": 0.25
}

Backend-Specific Load Override Notes

vLLM and vLLM Serve load override notes:

  • For vllm_serve, target_inflight maps to --max-num-seqs and also controls llm-pool admission.
  • vllm_max_model_len is the per-load context length.
  • vllm_kv_cache_dtype quantizes the KV cache; allowed UI values are auto, fp8, fp8_e4m3, fp8_e5m2. The service accepts any dtype string vLLM supports.
  • vllm_kv_cache_memory_bytes sets an absolute KV cache size in bytes. It is machine-independent and overrides vllm_gpu_memory_utilization for KV sizing. The load_constraints entry carries unit: "bytes" and display_unit: "mib" so the UI can present it in MiB.
  • Prefer keeping configured vllm_gpu_memory_utilization very low and controlling load-time cache budget with vllm_kv_cache_memory_bytes, otherwise vLLM may reserve most free VRAM.
  • vllm_max_pixels caps the vision-token budget per image for vision-language models; it is merged into the model's vllm_mm_processor_kwargs as max_pixels.
  • vllm_speculative_method selects the vLLM speculative path for this load, for example mtp, draft_model, or mlp_speculator. null disables the configured speculative path for that load.
  • vllm_speculative_model is the assistant/draft/speculator checkpoint or local path passed through vLLM's speculative_config.model. For Gemma 4 MTP this is the Gemma 4 assistant checkpoint, not a generic smaller draft model.
  • vllm_speculative_moe_backend maps to vLLM speculative_config.moe_backend. null clears the configured override for this load.
  • vllm_speculative_attention_backend maps to vLLM speculative_config.attention_backend. null clears the configured override for this load.
  • vllm_num_speculative_tokens maps to vLLM speculative_config.num_speculative_tokens.
  • For vllm_serve, target model id/path, binary path, library path, environment, host, port, API key, and extra CLI args are configured in the model definition, not overridden through the admin load body.
  • vllm_serve_extra_args must not set --max-num-seqs; target_inflight owns it.
  • Loading a vllm_serve model starts a local vllm serve subprocess. Unloading terminates that subprocess, so VRAM is released by the server process rather than by Python object cleanup alone.

TensorRT-LLM Serve load notes:

  • target_inflight maps to TensorRT-LLM max_batch_size and also controls llm-pool admission.
  • trtllm_max_seq_len, trtllm_max_num_tokens, and trtllm_enable_chunked_prefill map to the same top-level TensorRT-LLM YAML fields.
  • trtllm_kv_cache_memory_bytes maps to kv_cache_config.max_gpu_total_bytes. The API uses bytes; load_constraints tells clients to display MiB.
  • trtllm_kv_cache_dtype maps to kv_cache_config.dtype and accepts auto, fp8, or nvfp4.
  • llm-pool reads the configured base YAML, merges the effective values and the target_inflight batch size into a temporary YAML file, and passes that file through --config. The source YAML is unchanged. Normal runtime cleanup removes the temporary file.
  • TensorRT-LLM limits KV-cache memory to the lower result of max_gpu_total_bytes and free_gpu_memory_fraction. Keep the fractional value high enough to act only as a safety ceiling when an absolute budget should control allocation. This differs from the vLLM pattern of configuring a deliberately low utilization fraction alongside an absolute cache size.
  • Target model, binary, library path, environment, base YAML, parser names, host, port, timeouts, and extra CLI args remain in the model definition. Neither the base YAML nor trtllm_serve_extra_args may set max_batch_size.
  • Loading starts a local trtllm-serve process group. Unloading terminates that group, so VRAM is released by process exit.

SGLang Serve load notes:

  • target_inflight is passed to SGLang as --max-running-requests and also controls llm-pool admission. SGLang may reduce its native limit during KV-cache sizing; excess admitted requests then wait inside SGLang.
  • sglang_context_length, sglang_max_total_tokens, sglang_chunked_prefill_size, and sglang_kv_cache_dtype map to the corresponding SGLang server flags.
  • sglang_max_total_tokens sets the absolute KV-cache token capacity. sglang_mem_fraction_static remains a startup safety ceiling and must still be high enough to hold the target and assistant weights.
  • sglang_speculative_algorithm: "NEXTN" with a Gemma 4 assistant checkpoint selects SGLang's Frozen-KV MTP path. Set the algorithm to null to disable speculative decoding for one load.
  • The speculative step count and top-k map directly to SGLang server flags. When top-k is 1, llm-pool derives the draft-token count as the step count plus 1 and passes it explicitly; a conflicting explicit override is rejected.
  • sglang_serve_extra_args must not set --max-running-requests; target_inflight owns it.
  • Target model, binary, library path, environment, quantization, attention backend, host, port, parser names, timeouts, and extra CLI arguments remain in the model definition.
  • Loading starts a local sglang serve process group. Unloading terminates that group, so VRAM is released by process exit.

llama-server load override notes:

  • target_inflight maps to llama-server --parallel and also controls llm-pool admission.
  • llama_server_n_ctx maps to the native llama-server -c/--ctx-size flag for this load.
  • llama-server shares llama_server_n_ctx across its parallel slots. The approximate per-request context limit is llama_server_n_ctx / target_inflight.
  • llama_server_image_max_tokens maps to native --image-max-tokens and controls the per-image vision token budget.
  • llama_server_spec_type currently accepts only "draft-mtp" or null.
  • llama_server_spec_draft_n_max maps to native --spec-draft-n-max; the API constrains it to 1..6.
  • llama_server_spec_draft_p_min maps to native --spec-draft-p-min; the API constrains it to 0.0..1.0.
  • Model path, binary path, library path, mmproj, draft model path, GPU layers, flash attention, reasoning, host, port, API key, and extra native args are configured in the model definition, not overridden through the admin load body.
  • llama_server_extra_args must not set --parallel; target_inflight owns it.
  • Loading a llama_server model starts a local llama-server subprocess. Unloading terminates that subprocess, so VRAM is released by the native server process rather than by Python object cleanup alone.

exllama_cache_quant format:

  • omitted or null: fp16 KV cache
  • "<bits>": same quantization for K and V, for example "8"
  • "<k_bits>,<v_bits>": separate K/V quantization, for example "8,4"

exllama_cache_k_bits and exllama_cache_v_bits format:

  • both omitted: do not override the current configured value
  • both null: reset to fp16 KV cache
  • both integers from 2 through 8: override K and V separately
  • they must be provided together
  • they cannot be combined with exllama_cache_quant in the same load request

gguf_type_k and gguf_type_v format:

  • omitted or null: use the runtime default cache type
  • "<ggml_type_name>": a GGML cache type name, for example "f16", "q8_0", or "q4_0"

gguf_flash_attn format:

  • omitted: do not override the current configured value
  • "on": force Flash Attention on
  • "off": force Flash Attention off
  • "auto": use the runtime auto mode

Upstream references for these backend-specific value sets:

These overrides are runtime-only:

  • they do not modify settings.json
  • they do not modify local.json
  • they are returned separately from the configured definition in admin responses
  • replicas in the load request does not modify definition.replicas; it only selects the replica count for that one live load
  • target_inflight in the load request does not modify definition.target_inflight; it sets scheduler and managed-server concurrency for that one live load

Example response:

{
  "name": "google_gemma-4-E2B-it-Q8_0-gguf",
  "resolved_backend": "llama_cpp",
  "configured_enabled": false,
  "runtime_state": "loaded",
  "is_loaded": true,
  "replicas": 3,
  "replica_max": 4,
  "loaded_replicas": 3,
  "inflight_requests": 0,
  "queue_depth": 0,
  "runtime_inflight": 0,
  "configured_target_inflight": 1,
  "effective_target_inflight": 1,
  "fairness": {
    "rejected_per_key_limit": 0,
    "rejected_executor_limit": 0,
    "keys": []
  },
  "last_error": null,
  "vram_estimate_mib": 12340,
  "vram_estimate_replica_count": 3,
  "vram_estimate_source": "observed_load_delta",
  "load_constraints": {
    "gguf_n_ctx": {
      "kind": "integer",
      "minimum": 1,
      "step": 1
    },
    "gguf_flash_attn": {
      "kind": "enum",
      "default": "auto",
      "allowed_values": ["on", "off", "auto"],
      "examples": ["auto", "on", "off"]
    },
    "gguf_type_k": {
      "kind": "string_or_null",
      "format": "ggml_type_name",
      "default": "f16",
      "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
      "examples": ["f16", "q8_0", "q4_0"]
    },
    "gguf_type_v": {
      "kind": "string_or_null",
      "format": "ggml_type_name",
      "default": "f16",
      "allowed_values": ["f32", "f16", "bf16", "q8_0", "q4_0", "q4_1", "iq4_nl", "q5_0", "q5_1"],
      "examples": ["f16", "q8_0", "q4_0"]
    }
  },
  "load_recommendations": {
    "gguf_cache_type_pairs": {
      "kind": "pair_presets",
      "fields": ["gguf_type_k", "gguf_type_v"],
      "recommended_pairs": [
        {
          "label": "f16/f16",
          "gguf_type_k": "f16",
          "gguf_type_v": "f16"
        },
        {
          "label": "q8_0/q8_0",
          "gguf_type_k": "q8_0",
          "gguf_type_v": "q8_0"
        },
        {
          "label": "q4_0/q4_0",
          "gguf_type_k": "q4_0",
          "gguf_type_v": "q4_0"
        }
      ],
      "notes": [
        "Service-curated presets for GGUF cache types.",
        "Prefer symmetric GGUF K/V pairs by default; asymmetric pairs may reduce or disable GPU offload in upstream llama.cpp."
      ]
    }
  },
  "load_override": {
    "gguf_n_ctx": 32768,
    "gguf_flash_attn": "auto",
    "gguf_type_k": "q8_0",
    "gguf_type_v": "q4_0"
  },
  "capabilities": {
    "modalities": ["text"],
    "file_inputs": false,
    "multi_turn": true,
    "thinking_modes": ["default", "enabled", "disabled"],
    "response_formats": ["text"]
  },
  "definition": {
    "model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
    "backend": "llama_cpp",
    "prompt_format": "gemma4_template",
    "enable_thinking": null,
    "enabled": false,
    "replicas": 3,
    "replica_max": 4,
    "target_inflight": 1,
    "gguf_n_gpu_layers": -1,
    "gguf_n_ctx": 4096,
    "gguf_flash_attn": "auto",
    "gguf_type_k": null,
    "gguf_type_v": null
  }
}

Notes:

  • loading is allowed for configured models even when configured_enabled is false
  • after a successful load, vram_estimate_source may switch to observed_load_delta if a GPU delta could be measured during load
  • when replicas is provided, the model load is aggregate and all-or-nothing for that selected replica count
  • replicas and load overrides may only be changed while the model is unloaded or failed

Validation behavior:

  • 422 means the request body failed schema validation before runtime logic ran Examples: target_inflight: 0 gguf_n_ctx: 0 exllama_cache_size: 0 exllama_max_rq_tokens: 0 trtllm_kv_cache_memory_bytes: 0
  • 400 with code: "invalid_load_request" means the body was structurally valid, but the values were invalid for the resolved backend or runtime rules Examples: gguf_type_k: "q8-0" gguf_type_k: "foo" sending only exllama_cache_k_bits without exllama_cache_v_bits combining exllama_cache_quant with exllama_cache_k_bits/exllama_cache_v_bits exllama_cache_size: 8000 exllama_cache_quant: "fp16" llama_server_spec_type: "medusa" llama_server_spec_draft_n_max: 7 llama_server_spec_draft_p_min: 1.5 trtllm_kv_cache_dtype: "int8" sending ExLlamaV3-only fields to a llama_cpp model sending llama-server-only fields to a vLLM model sending load overrides while the model is already loaded and not first unloading it
  • 409 still applies for runtime state conflicts such as loading or unloading transitions Examples: loading a model while it is already unloading

POST /v1/admin/models/{model_name}/unload

Gracefully unloads one currently loaded model.

Rules:

  • 404 if model_name is unknown
  • 200 if already unloaded
  • 200 if already failed
  • 200 if already unloading
  • transition loaded -> unloading -> unloaded
  • unload stops all loaded replicas of the public model
  • new inference requests are rejected once unloading starts
  • in-flight requests are allowed to finish before resources are released

Selected fields from an unloaded response; the complete object has the same AdminModelEntry shape as an item returned by GET /v1/admin/models:

{
  "name": "google_gemma-4-E2B-it-Q8_0-gguf",
  "resolved_backend": "llama_cpp",
  "configured_enabled": false,
  "runtime_state": "unloaded",
  "is_loaded": false,
  "replicas": 3,
  "replica_max": 4,
  "loaded_replicas": 0,
  "inflight_requests": 0,
  "queue_depth": 0,
  "runtime_inflight": 0,
  "configured_target_inflight": 1,
  "effective_target_inflight": 1,
  "fairness": {
    "rejected_per_key_limit": 0,
    "rejected_executor_limit": 0,
    "keys": []
  },
  "last_error": null,
  "vram_estimate_mib": 12340,
  "vram_estimate_replica_count": 4,
  "vram_estimate_source": "observed_load_delta",
  "definition": {
    "model_path": "/home/gunnar/models/google_gemma-4-E2B-it-Q8_0/google_gemma-4-E2B-it-Q8_0.gguf",
    "backend": "llama_cpp",
    "device": "cuda",
    "prompt_format": "gemma4_template",
    "enabled": false,
    "replicas": 3,
    "replica_max": 4,
    "target_inflight": 1
  }
}

Unload And In-Flight Requests

Unload is graceful.

The service:

  • marks the model as unloading
  • rejects new inference requests for that model
  • cancels requests still waiting in the scheduler queue
  • waits for active backend requests to finish
  • releases runtime references and backend resources
  • marks the model as unloaded

inflight_requests tracks admitted requests so unload can wait safely while queued requests are cancelled and active requests drain.

Scheduler Alignment

The scheduler owns pending request queues. The runtime owns backend execution state.

On unload, the scheduler and runtime therefore:

  • stop new admissions
  • cancel queued work with model_unloading
  • let active backend work drain
  • unload the model after no admitted requests remain

Resource Release Guarantees

A successful unload means:

  • no runtime remains registered for the model
  • no new requests can reach that runtime
  • no in-flight requests remain
  • backend-owned objects are dereferenced
  • memory becomes reusable for later loads

The API does not promise that every allocator reports zero immediately after unload.

The guarantee is functional reuse, not cosmetic memory counters.

Documentation Expectations

The API supports UI clients through:

  • stable response models
  • OpenAPI descriptions on every admin endpoint
  • clear descriptions of runtime states
  • clear error codes for rejected inference requests

The UI should be able to render:

  • configured vs loaded state
  • current lifecycle state
  • last load error

Out Of Scope

The admin API does not define:

  • writes back to config files
  • force unload
  • active GPU job cancellation
  • retry policy for failed model loads