Repository navigation
Add /model_group/info endpoint for LiteLLM-aware clients - #196
Merged
JaeYeonLee0621 merged 3 commits intoOct 2, 2026
Merged
Conversation
Why: Clients such as oh-my-pi probe the standard LiteLLM management
endpoints (/model_group/info, /v2/model/info, /model/info,
/v1/model/info) before falling back to /v1/models. When none of them is
available, they derive the context window from the 'context_length'
field of /v1/models rows or from a bundled public model catalog matched
by model id, defaulting to 128k/32k when nothing matches. The gateway's
models are exposed under gateway-specific names (local-*, cloud-*),
which never match a public catalog, so such clients used the 128k
default for every model and mis-sized their context windows.
What: GET /model_group/info, authenticated with the same token auth as
the other gateway endpoints. It returns one JSON entry per configured
model:
- model_group / model_name: the public model name
- model_info: the router config's model_info passed through
(capabilities, costs, max_output_tokens), with max_input_tokens
derived from max_tokens (the context length) when not explicitly set
Example router config entry:
- model_name: local-qwen3.8-27b
litellm_params:
model: openai/Qwen/Qwen3.8-27B
api_base: https://vllm/v1
api_key: key
model_info:
id: local-qwen3.8-27b
mode: chat
supports_vision: true
supports_function_calling: true
supports_reasoning: true
supports_response_schema: true
supports_prompt_caching: true
max_tokens: 262144
input_cost_per_token: 0.0000004
output_cost_per_token: 0.0000003
cache_read_input_token_cost: 0.00000004
aliases:
- Qwen/Qwen3.8-27B
meffmadd
approved these changes
Sep 30, 2026
meffmadd
left a comment
Member
There was a problem hiding this comment.
LGTM! Thanks for the contribution! 🥳
Member
|
@Takalele Could you please fix the ruff formatting errors in CI? |
CI (ruff 0.16.9, `ruff check`) fails on test_endpoints.py:1225:9: I001 Import block is un-sorted or un-formatted: the inline import block in test_model_group_info_exposes_context_window listed `unittest.mock.patch` before `importlib.import_module`. Remove the inline `patch` import entirely — it is already imported at module level — leaving a single inline import that satisfies the check.
Contributor
Author
|
ci fix done see 4d3df72 |
CI's `ruff format --check` step fails on gateway/views/models.py: the /model_group/info change missed the two blank lines before the top-level definitions `_model_group_info_entry` (line 17) and the `model_group_info` view (line 67). Apply `ruff format` (0.16.9) to add them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add /model_group/info endpoint for LiteLLM-aware clients
Why: Clients such as oh-my-pi probe the standard LiteLLM management
endpoints (/model_group/info, /v2/model/info, /model/info,
/v1/model/info) before falling back to /v1/models. When none of them is
available, they derive the context window from the 'context_length'
field of /v1/models rows or from a bundled public model catalog matched
by model id, defaulting to 128k/32k when nothing matches. The gateway's
models are exposed under gateway-specific names (local-, cloud-),
which never match a public catalog, so such clients used the 128k
default for every model and mis-sized their context windows.
What: GET /model_group/info, authenticated with the same token auth as
the other gateway endpoints. It returns one JSON entry per configured
model:
(capabilities, costs, max_output_tokens), with max_input_tokens
derived from max_tokens (the context length) when not explicitly set
Example router config entry:
litellm_params:
model: openai/Qwen/Qwen3.8-27B
api_base: https://vllm/v1
api_key: <api_key>
model_info:
id: local-qwen3.8-27b
mode: chat
supports_vision: true
supports_function_calling: true
supports_reasoning: true
supports_response_schema: true
supports_prompt_caching: true
max_tokens: 262144
input_cost_per_token: 0.0000004
output_cost_per_token: 0.0000003
cache_read_input_token_cost: 0.00000004
aliases: