Skip to content

🥡 feat: Reload Custom Model Lists Across Replicas - #16471

Open
lia-by-librechat[bot] wants to merge 3 commits into
devfrom
lia/reload-model-catalog-dev
Open

lia-by-librechat[bot] wants to merge 3 commits into
devfrom
lia/reload-model-catalog-dev

Conversation

@lia-by-librechat

@lia-by-librechat lia-by-librechat Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Changing the default model list for an existing custom endpoint in librechat.yaml currently requires restarting every API replica. This PR validates a local or HTTP(S) CONFIG_PATH on an explicit admin reload, installs only existing custom endpoints' models.default changes, and reports every other YAML edit as requiring a restart. A new endpoint, changed credentials, memory policy, MCP settings, tool filters, and startup configuration never take effect through this reload.

The control appears in General settings only for a user with platform-scoped MANAGE_CONFIGS and ACCESS_ADMIN. It checks access when settings opens, not during every authenticated startup-config request. The POST independently enforces both grants. A model picker already open in another browser observes the applied replica generation and refreshes its models only after a serving replica proves it has that version.

Supersedes the closed broad-scope PR #16385 and replaces the canary-targeted PR #16470 with a clean dev-based change. Builds on last-good reload behavior merged in #16383. This PR targets dev so it can be merged before validating the live deployment on canary.

How it works

Admin POST → read and validate all YAML → report + project only existing custom models.default
           → install locally → compare-and-publish model-catalog:v1 generation in Redis
Other replica → background Redis check → reload its own source → install only on digest match
Open browser → poll authenticated local revision → fetch models with matching version header

A slower admin reload cannot overwrite a newer Redis generation. A replica whose source has not caught up retains its last good model list and retries at a configurable rate; publication is not a claim that all replicas have applied it. Only the process-local APP_CONFIG cache stores the accepted base. Redis is optional: without it, the operation and report are explicitly local-only. The initial Redis read is bounded; a server can start during a Redis outage and continue serving its last good config.

Rollout: configReload.clusterReady defaults to false. After all replicas on canary run this version, enable it in the source configuration and restart them before invoking reload. The new Redis key is versioned and isolated from the previous broad reload protocol. Do not perform clustered reloads during a mixed-version rollout. configReload.clientPollIntervalMs and configReload.mismatchRetryMs default to 3000 ms and 5000 ms respectively; remote fetches retain the previous 10-second timeout unless configured.

Type of change

  • Feature
  • Tests

Testing

  • Focused TypeScript backend tests for projection, generation fencing, lagging replicas, failures and authorization; focused Express route and client control/model-picker tests.
  • Real Redis and an HTTP config source with two independent Node API processes: one publishes, a lagging second keeps its old model list, then applies the new list after its source catches up; prefixed Redis key and invalid YAML retaining the last good config verified.
  • tsc --noEmit in packages/data-provider, packages/api, and client; builds of packages/data-provider, packages/data-schemas, packages/api, and packages/client; staged npm run static-checks (all affected gates passed).
  • Local Lighthouse was attempted but Vite could not resolve a missing cached @csstools/postcss-initial dependency in this worker; Chrome/Playwright is also not installed here. Full browser suites and canary deployment validation remain unrun. Lighthouse CI passed on this head; a full browser model-picker test remains unrun.

Risk / compatibility

Only default model lists for existing, uniquely named custom endpoints can change live. Other edits remain at each replica's startup value, including values used by the expired-file sweep, GitHub skill synchronization, MCP recovery, and global static tool catalog. Source-validation errors leave the last good base intact. Redis publication uses a single-key Lua compare-and-publish, and a published generation never exposes YAML or configuration secrets. MongoDB config override priority remains unchanged. An unavailable Redis store permits a local-only reload and reports propagation failure without blocking normal cached requests.

Checklist

  • Relevant regression tests added
  • Local focused checks passed
  • Lighthouse CI passed; browser model-picker and canary deployment tests remain pending

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Please review exact head d793891a7125d6160e47bb322351e1ec33dd7320. This replaces #16470 on current dev without canary-only commits. It limits live reload to existing custom endpoints' default model lists, gates cluster publication behind platform-managed activation and Redis CAS, and refreshes the browser's model picker after a replica applies the generation. Focused backend, HTTP, UI and three workspace typechecks pass. A real two-process HTTP/Redis test confirms publication, a lagging follower, recovery, and invalid-YAML retention. CI and local Lighthouse status are being collected.

@github-actions

Copy link
Copy Markdown
Contributor

Lighthouse CI failed. The last 80 log lines contain the measured budgets and assertion failures.

npm warn Unknown project config "allow-remote". This will stop working in the next major version of npm. See `npm help npmrc` for supported config options.

> LibreChat@v0.8.8-rc4 lighthouse:run
> playwright test --config=e2e/playwright.config.lighthouse.ts

[WebServer] [e2e] Started memory MongoDB at mongodb://127.0.0.1:32879/LibreChat-e2e

[WebServer] 2026-09-28 21:24:46 warn: [metrics] METRICS_SECRET is not set - /metrics will return 401 for all requests

[WebServer] 2026-09-28 21:24:46 info: Mongo Connection options

[WebServer] 2026-09-28 21:24:46 info: {
[WebServer]   "bufferCommands": false
[WebServer] }

[WebServer] 2026-09-28 21:24:46 info: Connected to MongoDB

[WebServer] 2026-09-28 21:24:46 warn: [Security] TRUST_TENANT_HEADER is active. Ensure your reverse proxy strips and sets X-Tenant-Id — untrusted clients must not be able to supply it directly.

[WebServer] 2026-09-28 21:24:52 info: [ensureBaseConfig] Loading base configuration...

[WebServer] 2026-09-28 21:24:52 info: Custom config file loaded:

[WebServer] 2026-09-28 21:24:52 info: {
[WebServer]   "version": "1.3.6",
[WebServer]   "cache": true,
[WebServer]   "endpoints": {
[WebServer]     "openAI": {
[WebServer]       "titleConvo": false
[WebServer]     }
[WebServer]   },
[WebServer]   "openapi": {
[WebServer]     "enabled": true
[WebServer]   }
[WebServer] }

[WebServer] 2026-09-28 21:24:52 info: [deploymentSkills] Loaded 1 deployment skill(s) from /home/runner/work/LibreChat/LibreChat/e2e/fixtures/deployment-skills

[WebServer] 2026-09-28 21:24:52 warn: The `CHECK_BALANCE` environment variable is deprecated. Please use the `balance` field in the `librechat.yaml` config file instead.
[WebServer] More info: https://librechat.ai/docs/configuration/librechat_yaml/object_structure/balance#overview

[WebServer] 2026-09-28 21:24:52 info: [credentials] New database detected. Credential fingerprints were recorded for future key-drift checks.

[WebServer] 2026-09-28 21:24:52 info: 
[WebServer] Outdated Config version: 1.3.6
[WebServer] Latest version: 1.3.17
[WebServer] 
[WebServer]       Check out the Config changelogs for the latest options and features added.
[WebServer] 
[WebServer]       https://www.librechat.ai/changelog
[WebServer] 
[WebServer] 

[WebServer] 2026-09-28 21:24:52 warn: RAG API is either not running or not reachable at http://127.0.0.1:8791, you may experience errors with file uploads.

[WebServer] 2026-09-28 21:24:53 error: Failed to start server: argument handler must be a function

Error: Process from config.webServer was not able to start. Exit code: 1

Open the full run

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Review exact pushed head a3b8edc4f220692a8bcb8a647b3787927e7eb7a5. This adds the dev-native model config middleware and a route-registration regression test, fixing the server-startup failure in Lighthouse and the three API shards on the previous head. Focused route tests and staged static checks passed. CI for this head is starting.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Please review exact pushed head ed83750f5d6f46e08acdd8184e57d0d0da14ab65. The final follow-up uses the repository-locked Tailwind 3 class order. CI on the preceding head passed Lighthouse and all API test shards; the sole static-check failure was the six class-order differences, corrected here. The focused reload-control test and staged static checks pass on this head. CI is rerunning.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review the latest head

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T01:28:54.786724Z ed83750 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ed83750f5d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +33 to +35
const next = revision.data?.generation;
if (authenticated && next != null && next > 0) {
setTarget((current) => Math.max(current ?? 0, next));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Handle generation resets instead of retaining the maximum

If the Redis generation key is lost or flushed, the Lua publisher restarts numbering at 1, and the backend tracker accepts that lower generation because it treats any unequal generation as a change. A browser that previously applied generation 5 will ignore the replacement generation 1 here due to Math.max, while applied.current also remains 5, so it never fetches the new catalog and stays stale until the counter exceeds 5 or the user session remounts. The revision protocol needs an epoch/opaque identifier or explicit reset handling rather than assuming the counter can never decrease.

Useful? React with 👍 / 👎.

Comment on lines +587 to +594
if (error instanceof ConfigReloadError) {
const validationErrors = error.validationErrors ?? [];
return res.status(400).json({
error:
validationErrors.length > 0
? 'Custom config validation failed'
: 'Custom config source could not be loaded',
validationErrors,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Sanitize validation issues before returning them

When the reloaded YAML fails validation, this sends complete raw Zod issues to the client. Those issues can contain submitted source material beyond the path, such as an invalid enum's received value, unrecognized key names, and custom parameter error text; this can disclose deployment configuration or accidentally embedded credentials to a config manager even though other config reads redact secrets. Map these failures to approved stable codes and safe metadata before serializing them, and let the UI localize those codes.

AGENTS.md reference: AGENTS.md:L122-L130

Useful? React with 👍 / 👎.

Comment on lines 21 to 23
const modelConfig = await loadModels(req);
res.set(CONFIG_GENERATION_HEADER, getConfigGenerationForConfig(req.config));
res.send(modelConfig);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Move generation stamping out of the legacy controller

Move the newly added generation lookup and response-header behavior into a TypeScript handler under packages/api, leaving this legacy CJS controller as wiring only. Calling the config-generation service and implementing the versioned response contract here extends backend behavior in /api, contrary to the repository's explicit module boundary and makes this reload protocol split across the legacy and typed layers.

AGENTS.md reference: AGENTS.md:L77-L81

Useful? React with 👍 / 👎.

Comment on lines 12 to +14
const { user } = useAuthContext();
const { data: startupConfig } = useGetStartupConfig();
const { data: configReloadAccess = false } = useConfigReloadAccessQuery(user?.id);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid probing admin access for every settings user

This query is enabled for every authenticated user who opens Settings, but its endpoint is mounted behind requireAdminAccess, whose normal denial path records a warning and performs a capability lookup. Consequently every non-admin settings visit generates an expected 403, misleading forbidden-access log traffic, and avoidable database work; because rejected query data is then treated as false, this happens again whenever the short-lived query is recreated or refetched. Use an authenticated, non-warning boolean capability probe or gate this request on an existing admin-access signal.

Useful? React with 👍 / 👎.

Comment on lines +408 to +415
async function getConfigRefreshStatus() {
const base = await ensureBaseConfig();
const generation = getConfigGenerationForConfig(base);
return {
distributed: syncConfigGeneration != null,
generation: generation === '' ? null : Number(generation),
pollIntervalMs:
base.config?.configReload?.clientPollIntervalMs ?? DEFAULT_CONFIG_RELOAD_CLIENT_POLL_MS,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Disable revision polling until clustered reload is enabled

When Redis is configured but configReload.clusterReady is absent or false, the reload endpoint rejects every clustered reload, yet this status still reports distributed: true. The client therefore polls /api/config/revision every three seconds for every visible authenticated session even though the feature is disabled by its default, adding potentially substantial request and Redis-read load to existing deployments. Gate distributed or expose a separate polling-enabled flag based on clusterReady so the new toggle's disabled default preserves prior behavior.

AGENTS.md reference: AGENTS.md:L90-L92

Useful? React with 👍 / 👎.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants