Skip to content

[Bug]: Bedrock GPT-6 Sol/Luna/Astra fall back to a 32k context, so long messages fail with empty_messages #16271

Description

@kimnamu

What happened?

Thanks for adding GPT-6 Sol and Luna in #16221 so quickly. On the Bedrock endpoint the GPT-6 profiles (us./global. openai.gpt-6-sol, -luna, -astra, and the gpt-5.6-* ones) match nothing in the Bedrock context map, so initializeAgent falls back to DEFAULT_MAX_CONTEXT_TOKENS (32,000). A single ~38k-token message is pruned away and the run fails with empty_messages before anything is sent. Called directly through Converse, Bedrock answered a 914k-token prompt on all three models.

Cause: openAIBedrockModels in packages/api/src/utils/tokens.ts only lists openai.gpt-oss-20b/120b. For Claude, bedrockModels reuses the native anthropicModels table, but nothing does that for the OpenAI entries.

Bedrock endpoint, Max Context Tokens empty Today (dev 59da55290) With the patch below
getModelMaxTokens(model, bedrock), us./global. Sol, Luna, Astra undefined 1,050,000 (the native gpt-6-* value)
budget from initializeAgent 30,400 997,500
38k-token message, same six IDs (live) ❌ empty_messages ✅ answered
openai.gpt-oss-120b-1:0 128,000 (budget 121,600) ✅ unchanged
Max Context Tokens set by hand used as-is ✅ used as-is (unchanged)

Workaround: set Max Context Tokens on the agent or conversation.

#14638 edits the same block (adds openai.gpt-5 / openai.gpt-5.5 at 272,000 for Mantle). Its keys don't match GPT-6 IDs, and for GPT-5.6 the longer key below wins.

I'd like to take this one: the patch and tests below are ready, and I'll open the PR once it's assigned.

Version Information

Source checkout: dev at 59da55290497347c5b76ed5bf192f9c036c35383, @librechat/agents 3.9.3 from the lockfile. Built-in Bedrock endpoint with BEDROCK_AWS_BEARER_TOKEN, BEDROCK_AWS_DEFAULT_REGION=us-east-1.

Steps to Reproduce

  1. Set BEDROCK_AWS_MODELS=us.openai.gpt-6-sol,us.openai.gpt-6-luna,us.openai.gpt-6-astra.
  2. Pick one of them, leave Max Context Tokens empty.
  3. Send one message of about 38k tokens (e.g. 13,000 words item0 item1 ... plus a question).
  4. The run fails with the context error instead of answering.

Reproduced with the jest probe below against live Bedrock, not in the browser.

Relevant log output

# packages/api jest probe: initializeBedrock -> createRun with the budget initializeAgent computes (live us-east-1)
us.openai.gpt-6-sol     | promptTokens(tiktoken)=38079 | getModelMaxTokens(bedrock)=undefined | maxContextTokens=30400 | ERR {"type":"empty_messages","info":"Message pruning removed all messages as none fit in the context window. ...
us.openai.gpt-6-luna    | promptTokens(tiktoken)=38079 | getModelMaxTokens(bedrock)=undefined | maxContextTokens=30400 | ERR {"type":"empty_messages", ...
us.openai.gpt-6-astra   | promptTokens(tiktoken)=38079 | getModelMaxTokens(bedrock)=undefined | maxContextTokens=30400 | ERR {"type":"empty_messages", ...
# global.openai.gpt-6-sol / -luna / -astra: same
Proposed patch, tests, and measured limits
 const openAIBedrockModels = {
   'openai.gpt-oss-20b': 128000,
   'openai.gpt-oss-120b': 128000,
+  /** Same models as the native OpenAI entries; Bedrock IDs carry the `openai.` prefix */
+  'openai.gpt-5.6': openAIModels['gpt-5.6'],
+  'openai.gpt-6-astra': openAIModels['gpt-6-astra'],
+  'openai.gpt-6-sol': openAIModels['gpt-6-sol'],
+  'openai.gpt-6-luna': openAIModels['gpt-6-luna'],
 };

New cases in packages/api/src/utils/tokens.spec.ts (us./global. Sol, Luna and Astra, us. GPT-5.6 Sol resolve to 1,050,000; gpt-oss stays 128,000). With the source change reverted they fail (Expected: 1050000, Received: undefined, 7 failed / 52 passed); with it, 59 passed.

Same probe on the patched branch:

us.openai.gpt-6-sol       | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items."
global.openai.gpt-6-sol   | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items."
us.openai.gpt-6-luna      | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items."
global.openai.gpt-6-luna  | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items listed."
us.openai.gpt-6-astra     | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items listed."
global.openai.gpt-6-astra | getModelMaxTokens(bedrock)=1050000 | maxContextTokens=997500 | OK | text="There are 13,000 items listed above."

Direct Converse calls (boto3, us-east-1, maxTokens 64): Sol, Luna and Astra each answered a 914,014-token prompt and returned context_length_exceeded for one of about 944k. So the native number is an upper bound on Bedrock: a prompt between ~944k and the 997,500 budget would still get a provider error. If you'd prefer a lower Bedrock-only value, I'm happy to change it.

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions