Skip to content

Add /model_group/info endpoint for LiteLLM-aware clients - #196

Merged
JaeYeonLee0621 merged 3 commits into
TU-Wien-dataLAB:mainfrom
Takalele:feature/model-group-info
Oct 2, 2026
Merged

JaeYeonLee0621 merged 3 commits into
TU-Wien-dataLAB:mainfrom
Takalele:feature/model-group-info

Conversation

@Takalele

Copy link
Copy Markdown
Contributor

Add /model_group/info endpoint for LiteLLM-aware clients
Why: Clients such as oh-my-pi probe the standard LiteLLM management
endpoints (/model_group/info, /v2/model/info, /model/info,
/v1/model/info) before falling back to /v1/models. When none of them is
available, they derive the context window from the 'context_length'
field of /v1/models rows or from a bundled public model catalog matched
by model id, defaulting to 128k/32k when nothing matches. The gateway's
models are exposed under gateway-specific names (local-, cloud-),
which never match a public catalog, so such clients used the 128k
default for every model and mis-sized their context windows.

What: GET /model_group/info, authenticated with the same token auth as
the other gateway endpoints. It returns one JSON entry per configured
model:

  • model_group / model_name: the public model name
  • model_info: the router config's model_info passed through
    (capabilities, costs, max_output_tokens), with max_input_tokens
    derived from max_tokens (the context length) when not explicitly set

Example router config entry:

  • model_name: local-qwen3.8-27b
    litellm_params:
    model: openai/Qwen/Qwen3.8-27B
    api_base: https://vllm/v1
    api_key: <api_key>
    model_info:
    id: local-qwen3.8-27b
    mode: chat
    supports_vision: true
    supports_function_calling: true
    supports_reasoning: true
    supports_response_schema: true
    supports_prompt_caching: true
    max_tokens: 262144
    input_cost_per_token: 0.0000004
    output_cost_per_token: 0.0000003
    cache_read_input_token_cost: 0.00000004
    aliases:
    • Qwen/Qwen3.8-27B
image

Why: Clients such as oh-my-pi probe the standard LiteLLM management
endpoints (/model_group/info, /v2/model/info, /model/info,
/v1/model/info) before falling back to /v1/models. When none of them is
available, they derive the context window from the 'context_length'
field of /v1/models rows or from a bundled public model catalog matched
by model id, defaulting to 128k/32k when nothing matches. The gateway's
models are exposed under gateway-specific names (local-*, cloud-*),
which never match a public catalog, so such clients used the 128k
default for every model and mis-sized their context windows.

What: GET /model_group/info, authenticated with the same token auth as
the other gateway endpoints. It returns one JSON entry per configured
model:
- model_group / model_name: the public model name
- model_info: the router config's model_info passed through
  (capabilities, costs, max_output_tokens), with max_input_tokens
  derived from max_tokens (the context length) when not explicitly set

Example router config entry:

- model_name: local-qwen3.8-27b
  litellm_params:
    model: openai/Qwen/Qwen3.8-27B
    api_base: https://vllm/v1
    api_key: key
  model_info:
    id: local-qwen3.8-27b
    mode: chat
    supports_vision: true
    supports_function_calling: true
    supports_reasoning: true
    supports_response_schema: true
    supports_prompt_caching: true
    max_tokens: 262144
    input_cost_per_token: 0.0000004
    output_cost_per_token: 0.0000003
    cache_read_input_token_cost: 0.00000004
    aliases:
    - Qwen/Qwen3.8-27B

@meffmadd meffmadd left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM! Thanks for the contribution! 🥳

@meffmadd

Copy link
Copy Markdown
Member

@Takalele Could you please fix the ruff formatting errors in CI?

CI (ruff 0.16.9, `ruff check`) fails on
test_endpoints.py:1225:9: I001 Import block is un-sorted or un-formatted:
the inline import block in test_model_group_info_exposes_context_window
listed `unittest.mock.patch` before `importlib.import_module`.

Remove the inline `patch` import entirely — it is already imported at
module level — leaving a single inline import that satisfies the check.
@Takalele

Takalele commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

ci fix done see 4d3df72

CI's `ruff format --check` step fails on gateway/views/models.py: the
/model_group/info change missed the two blank lines before the top-level
definitions `_model_group_info_entry` (line 17) and the
`model_group_info` view (line 67). Apply `ruff format` (0.16.9) to add
them.
@JaeYeonLee0621
JaeYeonLee0621 merged commit c92f86f into TU-Wien-dataLAB:main Oct 2, 2026
4 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants