Skip to content

Generated files skip delivery-path resolution, and non-PDF binaries still reach Chat Completions as file parts #16802

Description

@tommctech

Correction — 2026-10-08. The delivery claim in section 1 below is wrong, and is
superseded by this comment.
The storage half stands: a code output's stored text really is the HTML preview. But on
the native openAI route that text is not sent to the model — a follow-up turn adds
14 prompt tokens where the preview would need roughly 870. The rc3 400 is explained by
#16058 (landed in v0.8.8-rc4, four days after our rc3 image was built), not by a text
fallback replacing it. Section 2 is unaffected. Original text left intact below.

Split out of #16473 at @mihidumh's request, so the PR keeps its scope. Two gaps that need
a maintainer decision rather than a patch, with the measurements behind each.

Related: #16472, #16473.


1. A generated file skips delivery-path resolution

A file produced by the code interpreter (context: execute_code) does not get the
delivery-path resolution an upload gets. On v0.8.8-rc3 it therefore reached
formatDocumentBlock and was sent to api.openai.com/v1/chat/completions as

{ type: "file", file: { file_data: "data:text/csv;base64,…" } }

which that endpoint rejects for anything but PDF. The block persists in history, so the
conversation is dead from then on.

Repro, rc3 (v0.8.8-rc3, digest sha256:1eefa6a4c2b8…, agent on gpt-5.4-mini):
ask the agent to generate a small CSV, then send any follow-up message.

conversation 2213c7be-bc72-5479-a8a1-144b8d36b034
[0] user  textLen=111
[1] asst  blocks=tool_call,tool_call,text
          attachments=text/csv:top_10_ticket_openers_this_week.csv:252
[2] user  textLen=6                       <- "Thanks"
[3] asst  blocks=error

POST https://api.openai.com/v1/chat/completions -> 400 in 146ms (openai-processing-ms: 29)

anthropic is unaffected on the same build, because getAnthropicDocumentSource converts
an isAnthropicTextDocumentType into a {type: "text", media_type: "text/plain"} source.
The conversion that makes this safe already exists on one branch and not the other.

What changed on the released v0.8.8 — the 400 became a silent wrong answer

On librechat-ai/librechat-api:v0.8.8 (digest sha256:9cde97da2b50…) the same repro no
longer 400s, on openAI or on anthropic. resolveTurnLLMDeliveryPath resolves the
generated file to text, so it is sent as a text block and never reaches the document
path.

The conversation survives. The data does not.

The stored text for a generated file is its HTML preview

This is the part worth its own line. The text block carries the file's stored text, and
for a generated file that is a preview document, not the file:

acceptance.csv    type=text/csv    bytes=46
textFormat=html   textLen=3463
text starts: "<!DOCTYPE html>\n<html lang=\"en\">…<title>Preview</title>"

A 46-byte CSV is delivered to the model as ~3,463 characters of HTML markup. The model then answers about a file it has never seen, and nothing tells the user.

[Superseded — see the correction at the top. The file is stored that way, but is not delivered on the openAI route: measured at +14 prompt tokens.] Reproduced on
every generated file we have measured (46, 60, 65 and 101 bytes, all textFormat: html).

Two consequences:

  • A loud failure became a quiet one. The failed turn stopped the user; a confident answer
    about a file they watched the platform generate does not.
  • Roughly 75x token inflation per generated file, carried on every later turn.

Any fallback to stored text has this property, so it matters beyond this path.

One control still outstanding, stated for honesty: the rc3 failure was measured on
gpt-5.4-mini and the v0.8.8 passes on gpt-5.1. We have not yet run v0.8.8 with
gpt-5.4-mini. The block type is chosen before the request and both are plain OpenAI
models taking identical code paths, and OpenAI's rejection is endpoint-level validation of
the data URL's MIME type, so the model should not be able to account for the difference —
but we have not proved that, and will update here if it surprises us.


2. Non-PDF binaries still reach api.openai.com as file parts

Per @mihidumh on #16473, 99bc9973f covers the textual rows below; docx, xlsx and
zip are left as they are, because Azure OpenAI accepts docx as a file part and
rejecting every non-PDF binary for every OpenAI-shaped backend is a maintainer decision.
This is the evidence for that decision.

Setup: released v0.8.8, fileConfig.endpoints.agents.legacyFileUploadUX: true,
agent on gpt-5.1, native OpenAI — direct to api.openai.com/v1/chat/completions, no
gateway.
Each case is a new conversation with one file attached and a first message
asking about it. Fixtures are generated minimal valid examples of each type.

File Type Turn 1 Covered by 99bc9973f?
fixture.png image/png 200 n/a
fixture.pdf application/pdf 200, content read n/a
fixture.txt text/plain 400 yes
fixture.csv text/csv 400 yes
fixture.html text/html 400 yes
fixture.json application/json 400 yes
fixture.docx Word 400 no
fixture.xlsx Excel 400 no
fixture.zip application/zip 400 no

All seven failures carry the same message:

400 Invalid file data: 'messages[1].content[1].file.file_data'.
Expected a base64-encoded data URL with an application/pdf MIME type
(e.g. 'data:application/pdf;base64,SGVsbG8sIFdvcmxkIQ=='),
but got unsupported MIME type 'text/plain'.

All nine uploads were plain message attachments posted to /api/files with no
tool_resource, and every one was stored with llmDeliveryPath: "provider" — including
the two that succeeded. That matches resolveDefaultUploadLLMDeliveryPath, which returns
"provider" unconditionally when endpointConfig.legacyFileUploadUX === true, before any
MIME consideration. The flag alone is sufficient — no tool resource and no composer
choice — so this path is reproducible from a script.

On the default upload UX, same instance and build, all nine pass, and for the seven
carrying readable text the model demonstrably reads the content (each fixture embeds a
marker string the model echoes back).


The decision

Stating it as @mihidumh framed it, rather than proposing an answer:

  1. Whether generated files should be routed like uploads, through delivery-path
    resolution.
  2. Whether every non-PDF textual type should go as text on Chat Completions — partly
    addressed by 99bc9973f, opt-in via supportedMimeTypes for backends like Azure that
    accept them.
  3. Whether non-PDF binaries should be rejected for every OpenAI-shaped backend, given
    Azure accepts docx today and the code cannot distinguish the backends behind a
    gateway.

And, separately from all three: whether a fallback to stored text should send a preview
document.


Environment

librechat-ai/librechat-api:v0.8.8 (sha256:9cde97da2b50…) and
danny-avila/librechat-dev-api:latest at v0.8.8-rc3 (sha256:1eefa6a4c2b8…), Docker
Compose, Mongo 8.0.20. Native openAI and anthropic endpoints, no gateway. Measurements
come from an automated matrix, so re-running any of this on another route or type set is
cheap — happy to do so if it would help the decision.

Activity

  1. mihidumh commented on Oct 8, 2026

    @mihidumh
    Contributor

    @tommctech thanks for splitting this out. I measured section 1 on our deployment and read the path at tag v0.8.8. The storage half matches your report. The delivery half does not, and I think the code explains why.

    Setup. Released v0.8.8, fileConfig.endpoints.agents.legacyFileUploadUX: true, one custom OpenAI-compatible endpoint behind a LiteLLM proxy, Run Code on. Two conversations, one on a Claude route and one on a GPT route. Turn 1 asks the agent to write a 3-line CSV with Run Code. Turn 2 asks a question and forbids tool calls.

    The record is as you describe.

    probe_a.csv  type=text/csv  bytes=37  context=execute_code  source=local
    textFormat=html  textLen=3397  llmDeliveryPath=<unset>
    text starts: "<!DOCTYPE html>\n<html lang=\"en\">…"
    

    But the text is not sent. Input tokens per model call, as the gateway reports them from the provider (not LibreChat's own count):

    route last call of turn 1 turn 2 turn 3
    Claude 13,065 13,186 (+121) n/a
    GPT 7,715 (first call, same router tier as turn 2) 8,234 (+519: three tool calls with results, plus a 132-token question) 8,318 (+84)

    The preview page is about 1,000 tokens, and no delta has room for it. Both models also report no file content and no HTML in context.

    Why, at v0.8.8.

    • extractFileContext (packages/api/src/files/context.ts:73) adds a file's stored text only when llmDeliveryPath === 'text' or source === 'text'. A code output has neither: nothing under packages/api/src/files/code sets llmDeliveryPath, and the record's source is the storage strategy.
    • applyTurnDelivery (packages/api/src/agents/files/delivery.ts) does not give it a route either. It skips every record for which hasInferredLLMDeliveryPath is false, and that needs a stored path.
    • processAttachments in BaseClient.js drops the file from the media categories through isToolOwnedAttachment (packages/api/src/agents/attachments.ts:40), which is true for context === execute_code.

    So on v0.8.8 a generated file reaches the model through the sandbox only: not as a file part, and not as text.

    Your rc3 400. I think that is the bug #16058 fixed (db674eda60, 2026-09-18, first in v0.8.8-rc4, not in rc3). Its message: "a route-less record without a reference was classified as prompt content, so it counted toward the limit and could be encoded as media. Code outputs now stay tool-owned regardless of reference liveness". That would explain why the 400 went away between your two builds without a text fallback taking its place.

    What I did not check. The native openAI and anthropic endpoints you used, the child-agent run-file encoder (packages/api/src/agents/files/encode.ts), an agent handoff, and the Responses handlers. If your matrix can log usage.prompt_tokens for the turn before and the turn after the generated file, that would settle it on the native routes: a 46-byte CSV sent as its preview should add roughly 1,000 tokens.

    Your wider point stands either way: the stored text of a csv/docx/xlsx/pptx output is the HTML preview (getExtractedTextFormat, packages/api/src/files/code/extract.ts:101), so any future fallback that sends stored text for these records needs a textFormat check first.

  2. tommctech commented on Oct 8, 2026

    @tommctech
    Author

    Correction: you are right, and section 1's delivery claim is wrong. I measured it on
    the two native routes you listed as unchecked, and on openAI your result reproduces.

    I inferred delivery from reading the llmDeliveryPath === 'text' branch and never measured
    it. Worse, I had the disconfirming evidence in hand — I recorded llmDeliveryPath=<unset>
    on that same file record and read it as "the route does not expose this field" rather than
    as the condition being false. Apologies for the noise.

    Prompt tokens per model call

    From LibreChat's own transactions records (tokenType: "prompt"), so these are its
    accounting rather than gateway-reported provider usage — worth stating, since your figures
    came from the provider side.

    Native openAI, gpt-5.1, released v0.8.8:

    conversation call 1 call 2 turn 2
    generation 15,186 15,300 (+114) 15,314 (+14)
    baseline, no file 15,138 — 15,160 (+22)

    A follow-up adds 14 tokens. The stored preview for that file is 3,463 characters, roughly
    870 tokens, and there is no room for it. Your three code reasons hold on this route too:
    nothing sets llmDeliveryPath for a code output, applyTurnDelivery skips it, and
    isToolOwnedAttachment drops it from the media categories.

    The rc3 400 — your #16058 explanation fits exactly

    Our rc3 image was built 2026-09-14. db674eda60 landed 2026-09-18, first in
    v0.8.8-rc4. Our failing build predates the fix by four days, which accounts for the 400
    disappearing between our two builds without a text fallback replacing it. That also closes
    the control I flagged as outstanding — the difference was the build, not the model, so the
    gpt-5.4-mini rerun is no longer needed.

    One thing that does not fit, on the route you did not check

    Native anthropic, claude-sonnet-5, same build, same harness:

    conversation follow-up turn delta
    baseline, no file +24
    attach csv (upload) +36
    attach png (upload) +24
    generated file +852

    Full series for the generated case: 27,713 → 27,990 (+277, the call after the tool result)
    → 28,842 (+852, the follow-up).

    The turn-2 message is ~15 tokens and the turn-1 assistant reply was a few words, so roughly
    800 tokens appear on that route and not on the others. The stored preview for that file is
    3,463 characters ≈ 866 tokens, which is close enough to be worth your attention.

    Caveats, so this is not read as a mechanism claim: it is a token delta, not a captured
    request body; the counts are LibreChat's rather than the provider's; and a generation
    conversation has an extra model call, though the tool result is already inside call 2's
    prompt, so it should not land in the turn-2 delta. It could be the preview reaching the
    model on this path, or Anthropic-specific prompt accounting. I did not want to assert a
    mechanism twice in one issue.

    If it would help, I can capture the outbound request body for that turn, or run the same
    comparison across more file types — the matrix is scripted.

    What I still think stands

    Only the storage half, which you confirmed independently, plus the conclusion you drew from
    it: the stored text of a csv/docx/xlsx/pptx code output is the HTML preview
    (getExtractedTextFormat), so any future fallback that sends stored text for these records
    needs a textFormat check first.

    I have corrected the issue body above this comment rather than leaving the overstated
    version standing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    🐛 bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions