Skip to content

🏋️ fix: Bound Host File Edits Outside the API Event Loop - #16677

Merged
danny-avila merged 6 commits into
devfrom
lia/bound-host-edits
Oct 3, 2026
Merged

danny-avila merged 6 commits into
devfrom
lia/bound-host-edits

Conversation

@lia-by-librechat

@lia-by-librechat lia-by-librechat Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Exact replace_all batches introduced in #16520 can amplify tiny skill/sandbox files and repeatedly scan/rebuild large intermediate strings on the API event loop. At base 73d704656e39, the reported 54-edit sequence took 11,525 ms and prevented a scheduled timer from firing. The tolerant-window optimization in #16562 does not bound this exact-replacement workload.

Run host-side matching in a bounded reusable worker pool. Enforce cumulative scanned/reconstructed bytes, cumulative occurrence processing, edit count, intermediate output size, and a job deadline. Busy, over-budget, cancelled, or failed jobs write nothing. Attached-workspace edits retain their worker route and existing guarantees.

How It Works

native edit proposal → shared count/argument validation → existing file authorization/read
  → bounded worker admission, no waiting queue
  → existing matching strategies + per-batch work accounting
  → complete final result or safe rejection
  → cancellation/revision/permission/filter checks → one existing write

The pool admits two jobs per API process by default, reuses idle threads, and retires idle or cancelled workers. The job deadline uses a monotonic clock and is checked before accepting replies, independent of timer ordering. Deadline/cancellation termination releases admission only after the worker exits. No application services, storage, or credentials are loaded in the matcher worker. The existing matcher precedence, UTF-16 offsets, ambiguity handling, and ordered replacements are preserved. Admission bounds UTF-16 clone bytes; matching separately accounts UTF-8 input and each scan/reconstruction pass, reusing per-edit size measurements.

Configuration

Optional endpoints.agents.hostFileEdits controls host skill/non-attached-sandbox processing:

Field Default
maxEdits 100
maxWorkBytes 67,108,864
maxOccurrences 100,000
timeoutMs 2,000
maxConcurrent 2

Defaults apply when omitted. Occurrences count processing in each pass, not only unique replacements. Limits cannot be disabled; validated upper bounds remain in force. This intentionally rejects formerly unbounded work. Update all API replicas before enabling the new strict configuration field. No database migration or Code API change.

Verification

At head 96164f19dc865c2e646d013358eb48b98da9478e, 577 focused API tests, 371 configuration tests, and 29 legacy image-tool tests passed. API/data-provider typechecks, real data-provider/API builds, and touched-file static checks passed. Current-head independent review and CI status are recorded in the head comment.

Regressions cover compact amplification without timer starvation or writes, configured byte/occurrence limits, cancellation/deadlines and capacity recovery, final-write fences, pre-write/in-flight sandbox abort propagation, matching semantics, 9/10 MiB omitted-default compatibility with short and full matching context, multibyte text, full-context replace_all/contraction, actual 10 MiB sandbox persistence, deterministic work accounting, synchronous thread-creation failure, and safe errors through both skill and sandbox paths. Only host-controlled edit diagnostics are surfaced; unexpected failures use a generic message.

All five reported P2 findings are fixed. No findings were rejected. Independent review of 96164f19dc865c2e646d013358eb48b98da9478e found no new defects and passed 18 matcher probes, 2,000 differential cases, and three extracted sandbox-write probes. It remains incomplete: production worker/handler/configuration/persistence suites were unavailable in the isolated lane. Earlier incomplete reviews are not counted as clean reviews of this head. Exact-head CI Lighthouse, integrations, unit tests, builds, static checks and runtime smoke passed; all exact-head CI is now green (43 passed, 2 skipped) as of 2026-10-03 00:42 UTC.

Local Lighthouse is blocked by missing Chromium after a scratch-only Mongo socket workaround. Full local suites, config-migration/unused-package full gates, HTTP/model-provider exploit verification, and deployed-load benchmarks were not run. No UI layout changes.

Title verification: 🏋 has zero uses in 5,877 indexed LibreChat commits through 2026-10-01 21:39:51 UTC; sentinel passed. Frame: bounded CPU work moved off the API thread.

@lia-by-librechat

lia-by-librechat Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Head: 1f35df8ec6d3622d7f1b07a148ed41895ca54467

E adds cumulative edit-work limits and a bounded cancellable worker pool for host skill/sandbox edits. Attached edits retain their route.

  • 431 focused API tests and 369 configuration tests passed.
  • API/data-provider typechecks and real package builds passed.
  • Latest-dev baseline: 11,525 ms of timer starvation. Actual handler regressions reject the amplification batch without writing while timers run.
  • Static checks caught one unused test binding; its local cleanup passes and will be included in the next coherent push.
  • Local Lighthouse remains blocked by missing Chromium after a scratch-only Mongo socket workaround.
  • CI and exact-head independent review are running. No deployed exploit or load benchmark claimed.

@codegraph-librechat codegraph-librechat Bot added the 🗺️ Backend Platform codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) label Oct 2, 2026
@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Head: aa6a3d20c022193dd36a364af8cada5eaf692bf9

Host skill/sandbox edits have cumulative byte/occurrence limits, a server-side edit cap, and bounded off-thread processing. Attached edits retain their route.

  • 437 focused API tests, 369 configuration tests, and 29 legacy image-tool tests passed.
  • API/data-provider tsc --noEmit, real API build, and touched-file static checks passed.
  • E-R1 (P2) is fixed: monotonic deadlines reject late replies regardless of timer-callback order. Real-worker boundary and skill/sandbox handler regressions verify rejection, no writes, and capacity recovery.
  • The eager worker initialization and unused test binding behind first-head CI failures are fixed.
  • A fresh independent review and GitHub CI are running for this exact head.
  • Local Lighthouse could not launch Chromium after a scratch-only Mongo socket workaround. No deployed attack or load benchmark claimed.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Head: 4ce45079723dc98478eb3c1a5e61273fce086074
Base and merge-base: bb0543fd1519341a115036bec59c0ecd29228390

Synced with current dev without rewriting history. The configuration conflict retains both host-edit limits and subagent-activity policy. Security behavior is unchanged from the preceding head.

E-R1 (P2, late deadline acceptance) is fixed in aa6a3d20c022193dd36a364af8cada5eaf692bf9. The preceding head passed 952 focused tests, both changed-workspace typechecks, builds, static checks, and CI Lighthouse. Its independent source audit found no new defects but remained incomplete because isolated runtime verification could not run.

Current-head local verification and CI are running. A dedicated frozen-head review worktree is being prepared with its own dependencies so the next review can execute runtime checks. Earlier-head checks/review do not certify this synchronized head. No deployment or HTTP/model-provider exploit is claimed.

@lia-by-librechat

lia-by-librechat Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Head: 4ce45079723dc98478eb3c1a5e61273fce086074
Base and merge-base: bb0543fd1519341a115036bec59c0ecd29228390

E bounds host skill/sandbox edit work off the API event loop. Attached-worker edits retain their route.

  • 554 focused API tests, 371 configuration tests, and 29 legacy image-tool tests passed locally.
  • API/data-provider tsc --noEmit, real data-provider/data-schemas/API builds, and touched-file static checks passed.
  • CI has no failures. Lighthouse, both Redis integration lanes, builds, static checks, typechecks, API runtime smoke, and all backend test shards passed. Six memory E2E shards remain running as of 2026-10-02 17:41 UTC.
  • E-R1 (P2, late worker-reply acceptance) is fixed in aa6a3d20c022193dd36a364af8cada5eaf692bf9, an ancestor of this head. Deadline and handler no-write regressions pass. No rejected findings.
  • Independent review of this exact head is incomplete, with no new source findings. It audited all changed files and confirmed E-R1 in source, but could not run its isolated runtime checks or independently recompute the merge-base because workspace admission was unavailable. This is not a clean runtime review.
  • Local Lighthouse could not launch missing Chromium; CI Lighthouse passed. Complete local suites, config-migration/unused-package full gates, a live HTTP/model-provider exploit, and a deployed-load benchmark were not run.

No merge or deployment performed.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review the latest head, final review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T19:19:11.668925Z 0f283de Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4ce4507972

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/data-provider/src/config.ts Outdated
.int()
.min(1024)
.max(256 * 1024 * 1024)
.default(32 * 1024 * 1024),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve edits for files within the authoring limit

When hostFileEdits is omitted, a single exact edit of an ASCII file larger than roughly 8 MiB but within the existing 10 MiB authoring limit is rejected: the matcher charges the working file once before matching, once during exact matching, and roughly twice during reconstruction, so a 9 MiB file consumes about 36 MiB against this 32 MiB default. Such files were accepted previously and still satisfy MAX_AUTHORING_BYTES; raise the default or adjust the accounting so the ordinary single-edit path remains supported.

AGENTS.md reference: AGENTS.md:L90-L92

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d97c7ca5d2b949fdb4ec757d79327a3380870436. The default cumulative work budget is 64 MiB, preserving ordinary exact edits through the existing 10 MiB authoring limit. Occurrence, batch-count, output-size, deadline, and concurrency bounds remain enforced. New real-worker 9/10 MiB regressions and a real 10 MiB sandbox-handler edit pass; amplification and configured lower-budget rejection tests still pass.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The compatibility sweep found a residual case when old_text spans the full file. Fixed in 0f283deda29175e02fbb499a2fa09c40f5863d67: admission bounds actual UTF-16 clone units, and worker accounting reuses each edit's byte measurements without redundant scan charges. Default remains 64 MiB. Full 9/10 MiB ASCII and 10 MiB UTF-8 contexts, full-context replace_all/contraction, and real sandbox persistence pass; cumulative work limits and amplification rejection remain green.

Comment thread packages/api/src/agents/handlers.ts Outdated
Comment on lines 4233 to 4235
} catch (error) {
if (signal?.aborted) throw error;
return errorResult(tc, error instanceof Error ? error.message : 'Failed to edit file');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Sanitize synchronous worker startup failures

If worker initialization throws synchronously—for example because thread creation fails under resource pressure—this catch returns the raw operational error.message as the tool result; the equivalent skill-file path does the same. Unlike HostEditError, such messages are not controlled and may expose absolute paths or runtime diagnostics to a persisted, user-visible surface, so only known safe edit errors should be forwarded and all other failures should use the existing generic message.

AGENTS.md reference: AGENTS.md:L122-L128

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d97c7ca5d2b949fdb4ec757d79327a3380870436. Only host-controlled HostEditError diagnostics are forwarded. Skill and sandbox handlers now use a generic message for unexpected failures, and synchronous worker acquisition failures are sanitized before admission is charged. Regressions cover constructor failure/capacity recovery, both skill write paths, synchronous/asynchronous sandbox failures, no diagnostic disclosure, and no writes. Safe edit feedback remains unchanged.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Head: d97c7ca5d2b949fdb4ec757d79327a3380870436

Both Codex P2 findings are fixed: the default work budget supports ordinary exact edits through the 10 MiB authoring limit, and uncontrolled worker startup failures cannot reach tool results. A boundary sweep rechecked admission, budgets, cancellation, deadlines, controlled diagnostics, and final-write fences.

  • 563 focused API tests, 371 configuration tests, and 29 legacy image-tool tests passed.
  • API/data-provider tsc --noEmit, real data-provider/API builds, and touched-file static checks passed.
  • New regressions cover 9/10 MiB omitted-default compatibility, real sandbox persistence, synchronous constructor failure with capacity recovery, both skill paths, safe sync/async sandbox failures, and no writes on failure.
  • A fresh independent review and CI cover this exact head. Earlier incomplete reviews are not counted as clean.
  • E-R1 remains fixed in ancestor aa6a3d20c022193dd36a364af8cada5eaf692bf9. Local Lighthouse remains blocked by missing Chromium. No merge/deployment or live HTTP/model-provider exploit claimed.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review the latest head, final review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d97c7ca5d2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/api/src/agents/handlers.ts Outdated
const ctx = sandboxSessionContext(tc, sandboxContext);
let writeResult: Awaited<ReturnType<NonNullable<ToolExecuteOptions['writeSandboxFile']>>>;
try {
signal?.throwIfAborted();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Let pre-write cancellation escape the sandbox catch

When the run signal aborts after host edit processing completes but before the sandbox write starts, this check throws an AbortError inside the surrounding try, so the catch immediately converts it into an ordinary tool error. The run cancellation machinery therefore cannot recognize the cancellation; move the check outside this catch or rethrow abort errors before translating write failures.

AGENTS.md reference: AGENTS.md:L45-L50

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 96164f19dc865c2e646d013358eb48b98da9478e. The final pre-write abort check is outside the write-failure catch; an in-flight write AbortError with an aborted run signal is rethrown before error translation. Both real-handler regressions fail on the previous head and pass now: pre-write cancellation dispatches no write, and both cases reach the run-abort classification without a success artifact or ordinary write-error warning. Normal write failures retain their existing behavior.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Head: 0f283deda29175e02fbb499a2fa09c40f5863d67
Merge-base: bb0543fd1519341a115036bec59c0ecd29228390

The residual full-context P2 finding is fixed. Admission bounds UTF-16 clone bytes; the worker separately accounts UTF-8 input, matching and reconstruction. Per-edit byte measurements are reused rather than recharging repeated measurements. The 64 MiB default is unchanged. Configured lower-budget and amplification rejection remain enforced.

  • 575 focused API tests, 371 configuration tests, and 29 legacy image-tool tests passed.
  • API/data-provider typechecks, real API build, and touched-file static checks passed.
  • New coverage includes full 9/10 MiB ASCII context, 10 MiB UTF-8 context, full-context replace_all, contraction, real sandbox persistence, clone admission limits, deterministic cumulative accounting, and UTF-8 size transitions.
  • E-R1 (deadline acceptance), E-C1 (file-size compatibility), E-C2 (startup disclosure), and E-R2 (large matching context) are fixed. No findings rejected.
  • Current-head independent review and CI are running. Earlier incomplete reviews are not reported as clean. Local Lighthouse remains unavailable; no deployment or HTTP/model-provider exploit claimed.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review the latest head, final review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Delightful!

Reviewed commit: 0f283deda2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@lia-by-librechat

lia-by-librechat Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Head: 96164f19dc865c2e646d013358eb48b98da9478e

Exact-head CI is fully green: 43 passed, 2 skipped, no failed/pending checks. All five reported P2 findings have fixing commits and regression coverage recorded above.

The user-requested independent runtime retry is blocked before execution. Git repaired the mandatory reviewer-owned linked worktree, but the worker still rejects its cwd as “not a linked worktree” (HTTP 422). Reviewer policy requires that exact lane, so no substitute was used. No production-runtime checks ran in this retry and the independent review remains incomplete. The earlier source audit, 18 matcher probes, 2,000 differential cases and three write probes remain the verified independent evidence.

No source changes, merge, or deployment. Restore review-lane routing before another runtime attempt.

@danny-avila
danny-avila merged commit 4217b98 into dev Oct 3, 2026
45 checks passed
@danny-avila
danny-avila deleted the lia/bound-host-edits branch October 3, 2026 00:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🗺️ Backend Platform codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants