Skip to content

📨 fix: Resolve Stored Responses by Their IDs - #16456

Open
lia-by-librechat[bot] wants to merge 2 commits into
canaryfrom
lia/responses-stored-id-canary
Open

lia-by-librechat[bot] wants to merge 2 commits into
canaryfrom
lia/responses-stored-id-canary

Conversation

@lia-by-librechat

Copy link
Copy Markdown
Contributor

What breaks

POST /api/agents/v1/responses returns resp_… but GET /responses/:id and previous_response_id looked for a conversation with that ID. The reply is actually a message in a separate conversation, so clients cannot retrieve or continue the response they were handed. store: true also passed Express req where saveMessage expects an authenticated { userId } context, and streamed response.completed was emitted before any required write. Save failures were logged and ignored.

What changes

  • Resolve resp_… via the existing indexed, owner-scoped getMessage, then verify the owner-scoped conversation. Retain authorized conversation-ID reads for existing clients. A returned response ID retrieves only its own reply; continuation replays only through that reply.
  • Persist user input, then a pending assistant reply, then the conversation snapshot. Mark the assistant reply stop only once the snapshot exists. Pending or failed writes cannot be retrieved or reused, even on an existing conversation. Read errors fail rather than pretending the history was empty. No automatic retries.
  • For streamed responses, close and materialize output before saving, then emit response.completed only after the writes commit. Storage errors emit response.failed with a bounded error and [DONE]. JSON errors return 500 instead of a successful body. Reply announcements remain best-effort after durable storage.
owner + resp_id -> indexed assistant message -> owner-scoped conversation
store:true -> input -> pending reply -> conversation snapshot -> committed reply -> completed response

Related to AI-1593 and #14466.

Verification

  • Controller regressions: two-turn POST → GET → POST, SSE and JSON, owner denial, unstored IDs, missing and failed history, input/conversation/commit write failures, ordered completion.
  • Real mongodb-memory-server contract test exercises saveMessage, saveConvo, updateMessage, getMessage, and getConvo with pending/committed/other-owner responses.
  • Responses handlers and client-tool specs, packages/api typecheck, package build and scoped static checks.

Limits

This PR does not change tool execution, subagent discovery, typed history storage, reasoning replay, or Anthropic protocol support. Older conversation-ID requests keep their previous behavior. Pending partial rows are not automatically retried or declared complete. The separate subagent parity slice can land independently; coordinate if it starts changing responses.js in the same release window.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Review handoff for exact head 115f2f677d7faf95ae13048456cee51fa6bd2a4d: resp_ IDs now resolve via the owner-scoped persisted assistant message and its conversation. Both response modes require input + reply + conversation + final reply marker before reporting completion. Please check GET/POST ownership, partial-write/stream-failure behavior, and the real MongoDB contract test. The subagent parity slice is independent.

@github-actions

Copy link
Copy Markdown
Contributor

Lighthouse CI failed. The last 80 log lines contain the measured budgets and assertion failures.

│ 23      │ 'http://localhost:3080/api/permissions/mcpServer/effective/all'                                                 │ 2988.4890000000014 │ 3749.2700000000186 │ 200    │
│ 24      │ 'http://localhost:3080/api/prompts/groups?limit=10'                                                             │ 2988.695000000007  │ 4253.917000000016  │ 200    │
│ 25      │ 'http://localhost:3080/api/keys?name=openAI'                                                                    │ 3270.6290000000736 │ 3906.8310000000056 │ 200    │
│ 26      │ 'http://localhost:3080/api/presets'                                                                             │ 3270.951000000059  │ 3924.6230000000214 │ 200    │
│ 27      │ 'http://localhost:3080/api/tags'                                                                                │ 3271.1560000000172 │ 3926.9980000000214 │ 200    │
│ 28      │ 'http://localhost:3080/api/share/link/16390000-0000-4000-8000-000000000001'                                     │ 3272.896000000008  │ 4255.506000000052  │ 200    │
│ 29      │ 'http://localhost:3080/api/messages/16390000-0000-4000-8000-000000000001'                                       │ 3273.103000000003  │ 4413.435999999987  │ 200    │
│ 30      │ 'http://localhost:3080/api/files/config'                                                                        │ 3274.0780000000377 │ 4179.4310000000405 │ 200    │
│ 31      │ 'http://localhost:3080/api/user/settings/favorites/tools'                                                       │ 3275.1900000000023 │ 4432.125000000058  │ 200    │
│ 32      │ 'http://localhost:3080/api/endpoints/token-config'                                                              │ 3275.3820000000414 │ 4437.5470000000205 │ 200    │
│ 33      │ 'http://localhost:3080/api/user/settings/skills/active'                                                         │ 3275.560000000056  │ 4762.484000000055  │ 200    │
│ 34      │ 'http://localhost:3080/api/agents/tools/web_search/auth'                                                        │ 3275.725000000035  │ 7276.58600000001   │ 200    │
│ 35      │ 'http://localhost:3080/api/agents/tools/calls?conversationId=16390000-0000-4000-8000-000000000001'              │ 3276.8480000000563 │ 4763.781000000017  │ 200    │
│ 36      │ 'http://localhost:3080/api/agents/chat/status/16390000-0000-4000-8000-000000000001?generationProtocolVersion=2' │ 4522.167000000016  │ 4775.483999999997  │ 200    │
└─────────┴─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴────────────────────┴────────────────────┴────────┘

Inspect .lighthouse HTML/JSON and e2e/lighthouse/README.md. Reuse loaded user/config data; overlap independent reads without bypassing authorization.

┌─────────┬────────────────────────────┬──────────────────────┬───────┐
│ (index) │ audit                      │ median               │ limit │
├─────────┼────────────────────────────┼──────────────────────┼───────┤
│ 0       │ 'largest-contentful-paint' │ 4537.764             │ 4500  │
│ 1       │ 'cumulative-layout-shift'  │ 0.017903602683652913 │ 0.1   │
│ 2       │ 'total-blocking-time'      │ 293.4549999999999    │ 500   │
└─────────┴────────────────────────────┴──────────────────────┴───────┘

  1) [chrome] › e2e/lighthouse/load.spec.ts:10:5 › serial database latency stays within web-vitals budgets 

    Error: Median largest-contentful-paint must stay within 4500

    expect(received).toBeLessThanOrEqual(expected)

    Expected: <= 4500
    Received:    4537.764

       at audit.ts:159

      157 |   console.table(measured);
      158 |   for (const { audit, median, limit } of measured) {
    > 159 |     expect(median, `Median ${audit} must stay within ${limit}`).toBeLessThanOrEqual(limit);
          |                                                                 ^
      160 |   }
      161 |   return results;
      162 | }
        at auditPage (/home/runner/work/LibreChat/LibreChat/e2e/lighthouse/audit.ts:159:65)
        at /home/runner/work/LibreChat/LibreChat/e2e/lighthouse/load.spec.ts:33:19

    attachment #1: screenshot (image/png) ──────────────────────────────────────────────────────────
    e2e/lighthouse/.test-results/load-serial-database-latency-stays-within-web-vitals-budgets-chrome/test-failed-1.png
    ────────────────────────────────────────────────────────────────────────────────────────────────

    Error Context: e2e/lighthouse/.test-results/load-serial-database-latency-stays-within-web-vitals-budgets-chrome/error-context.md

    attachment #3: trace (application/zip) ─────────────────────────────────────────────────────────
    e2e/lighthouse/.test-results/load-serial-database-latency-stays-within-web-vitals-budgets-chrome/trace.zip
    Usage:

        npx playwright show-trace e2e/lighthouse/.test-results/load-serial-database-latency-stays-within-web-vitals-budgets-chrome/trace.zip

    ────────────────────────────────────────────────────────────────────────────────────────────────


🤖: global teardown has been started
2026-09-28 16:36:41 �[32minfo�[39m: �[32mMongo Connection options�[39m
2026-09-28 16:36:41 �[32minfo�[39m: �[32m{�[39m
�[32m  "bufferCommands": false�[39m
�[32m}�[39m
🤖:  ✅  Connected to Database
🤖:  ✅  Found user in Database
🤖:  ✅  Deleted 1 convos & 2 messages
🤖:  ✅  Deleted user from Database
🤖: global teardown has been started
2026-09-28 16:36:42 �[32minfo�[39m: �[32mMongo Connection options�[39m
2026-09-28 16:36:42 �[32minfo�[39m: �[32m{�[39m
�[32m  "bufferCommands": false�[39m
�[32m}�[39m
🤖:  ✅  Connected to Database
🤖:  ⚠️  User not found in Database
  1 failed
    [chrome] › e2e/lighthouse/load.spec.ts:10:5 › serial database latency stays within web-vitals budgets 

Open the full run

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Review handoff for exact pushed head 385b15018af7c523b519802c57465cab9c8c9e08: the owner-scoped response-ID resolver and committed-reply marker are unchanged; this head adds a real stream-emitter test proving final text is sealed without response.completed until the host explicitly completes it. Controller tests cover SSE/JSON two-turn retrieval, other-user denial, unstored IDs, write failures and history-read errors; real MongoDB test covers the indexed lookup. Please review this head.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-28T17:36:41.492121Z 385b150 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 385b15018a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

? messages.findIndex((message) => message.messageId === responseId)
: -1;
if (responseId && lastIndex < 0) return [];
return (responseId ? messages.slice(0, lastIndex + 1) : messages).filter(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reconstruct the response ancestry instead of slicing history

When a client branches from an older response, this chronological slice includes unrelated sibling turns. For example, after resp_1 → resp_2, creating resp_3 from resp_1 stores all three in one conversation; subsequently continuing resp_3 slices through resp_2, so the model receives the sibling user prompt and answer even though they are not ancestors of resp_3. Persist and traverse the previous_response_id relationship, or otherwise isolate branches, rather than treating every earlier conversation message as history.

Useful? React with 👍 / 👎.

Comment on lines +667 to +671
previousResponse = await resolveStoredResponse(
principal.userId,
request.previous_response_id,
db,
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Move response-resolution orchestration out of /api

Move this new validation, database resolution, error mapping, and branching into a TypeScript service in packages/api, leaving this CJS controller to pass request data and invoke it. Keeping the behavior here directly violates the repository boundary for edits to existing /api files and makes the new storage protocol harder to reuse and typecheck.

AGENTS.md reference: AGENTS.md:L64-L68

Useful? React with 👍 / 👎.

Comment on lines +773 to +778
const previousMessages = previousResponse
? await loadPreviousMessages(
conversationId,
principal.userId,
previousResponse.message?.messageId,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Avoid serial reads when loading response history

For a resp_… continuation, resolveStoredResponse serially awaits getMessage and getConvo, and this block then performs getMessages only after both have completed (with admission reads also occurring in between). Thus every continued response now pays at least three sequential Mongo round trips; return or load the conversation and history together, or parallelize the latter reads once the message identifies its conversation. The repository specifically requires message-loading changes to avoid serial database reads, and its CI models 250 ms per query.

AGENTS.md reference: AGENTS.md:L51-L54

Useful? React with 👍 / 👎.

Comment on lines +30 to +32
return (responseId ? messages.slice(0, lastIndex + 1) : messages).filter(
(message) =>
!message.messageId?.startsWith(RESPONSE_ID_PREFIX) || message.finish_reason === 'stop',

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exclude the entire turn when its storage commit fails

If saveConversation or the final updateMessage fails after the input was saved, the request reports response_storage_failed, but this filter removes only the pending_storage assistant row and retains the associated user messages. A later continuation via the supported legacy conversation ID—and any later response-ID continuation created after that—therefore sends the supposedly failed prompt to the model. The commit protocol needs to roll back the input or tag every message in the turn so history selection can exclude the whole uncommitted turn.

Useful? React with 👍 / 👎.


// Load messages for this conversation
const messages = await db.getMessages({ conversationId: responseId, user: userId });
const { conversation, message } = stored;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Return the stored agent ID as the response model

For the newly supported response-ID lookup, the owning conversation stores the requested agent under agent_id, while the response builder later reads conversation.agentId || conversation.model. Lean conversation documents do not alias agent_id to agentId, so a retrieved response reports the provider model (or unknown) instead of the agent ID that the original POST returned in model. Use conversation.agent_id or persist the original response model with the message.

Useful? React with 👍 / 👎.


// Load messages for this conversation
const messages = await db.getMessages({ conversationId: responseId, user: userId });
const { conversation, message } = stored;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use response timestamps when retrieving by response ID

This response-ID path has the specific message, but the returned created_at and completed_at are still derived from the conversation. After later continuations, retrieving an earlier response therefore reports the conversation's original creation time and its latest update time, so the same stored response's completion timestamp changes whenever another turn is added. Use the resolved message's timestamps for response-ID reads; the new StoredResponseMessage contract already exposes createdAt and updatedAt.

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants