Bedrock GPT-6 Sol/Luna/Astra fall back to a 32k context, so long messages fail with empty_messages #16272
Replies: 2 comments
This comment was marked as spam.
This comment was marked as spam.
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What happened?
Thanks for adding GPT-6 Sol and Luna in #16221 so quickly. On the Bedrock endpoint the GPT-6 profiles (
us./global.openai.gpt-6-sol,-luna,-astra, and thegpt-5.6-*ones) match nothing in the Bedrock context map, soinitializeAgentfalls back toDEFAULT_MAX_CONTEXT_TOKENS(32,000). A single ~38k-token message is pruned away and the run fails withempty_messagesbefore anything is sent. Called directly through Converse, Bedrock answered a 914k-token prompt on all three models.Cause:
openAIBedrockModelsinpackages/api/src/utils/tokens.tsonly listsopenai.gpt-oss-20b/120b. For Claude,bedrockModelsreuses the nativeanthropicModelstable, but nothing does that for the OpenAI entries.dev59da55290)getModelMaxTokens(model, bedrock), us./global. Sol, Luna, Astraundefinedgpt-6-*value)initializeAgentempty_messagesopenai.gpt-oss-120b-1:0Workaround: set Max Context Tokens on the agent or conversation.
#14638 edits the same block (adds
openai.gpt-5/openai.gpt-5.5at 272,000 for Mantle). Its keys don't match GPT-6 IDs, and for GPT-5.6 the longer key below wins.I'd like to take this one: the patch and tests below are ready, and I'll open the PR once it's assigned.
Version Information
Source checkout:
devat59da55290497347c5b76ed5bf192f9c036c35383,@librechat/agents3.9.3 from the lockfile. Built-in Bedrock endpoint withBEDROCK_AWS_BEARER_TOKEN,BEDROCK_AWS_DEFAULT_REGION=us-east-1.Steps to Reproduce
BEDROCK_AWS_MODELS=us.openai.gpt-6-sol,us.openai.gpt-6-luna,us.openai.gpt-6-astra.item0 item1 ...plus a question).Reproduced with the jest probe below against live Bedrock, not in the browser.
Relevant log output
Proposed patch, tests, and measured limits
const openAIBedrockModels = { 'openai.gpt-oss-20b': 128000, 'openai.gpt-oss-120b': 128000, + /** Same models as the native OpenAI entries; Bedrock IDs carry the `openai.` prefix */ + 'openai.gpt-5.6': openAIModels['gpt-5.6'], + 'openai.gpt-6-astra': openAIModels['gpt-6-astra'], + 'openai.gpt-6-sol': openAIModels['gpt-6-sol'], + 'openai.gpt-6-luna': openAIModels['gpt-6-luna'], };New cases in
packages/api/src/utils/tokens.spec.ts(us./global. Sol, Luna and Astra, us. GPT-5.6 Sol resolve to 1,050,000; gpt-oss stays 128,000). With the source change reverted they fail (Expected: 1050000, Received: undefined, 7 failed / 52 passed); with it, 59 passed.Same probe on the patched branch:
Direct Converse calls (boto3, us-east-1, maxTokens 64): Sol, Luna and Astra each answered a 914,014-token prompt and returned
context_length_exceededfor one of about 944k. So the native number is an upper bound on Bedrock: a prompt between ~944k and the 997,500 budget would still get a provider error. If you'd prefer a lower Bedrock-only value, I'm happy to change it.Code of Conduct
All reactions