Skip to content

fix(memory): persist smaller windows after extraction timeouts - #1769

Merged
Teingi merged 2 commits into
oceanbase:masterfrom
knqiufan:codex/fix-memory-window-timeout-1757
Oct 1, 2026
Merged

Teingi merged 2 commits into
oceanbase:masterfrom
knqiufan:codex/fix-memory-window-timeout-1757

Conversation

@knqiufan

Copy link
Copy Markdown
Contributor

Which issue or RFC does this PR close?

Closes #1757. Updates RFC 1515's Memory recovery behavior.

Rationale for this change

When a large Memory extraction times out, retrying the same Source window can indefinitely block subsequent input in that Scope. Other Scopes can continue and manual recovery remains possible, but retries currently do not adapt the failing input size.

What changes are included in this PR?

  • On a generation timeout, halve the failed journal window (minimum one position) for the next invocation. Retain the reduction in a Memory-owned per-Scope table so new Workers and Server restarts do not retry the original large window.
  • Preserve the failed Source cursor position and leave the invocation unacknowledged. Updating the recovery hint requires the current processing fence and the cursor CAS; a stale timeout cannot overwrite concurrently committed progress.
  • Retain the smaller limit through the backlog observed at failure, then clear it atomically with successful Memory publication, cursor advancement, and invocation acknowledgement. Configured and explicit flush limits still bound each invocation.
  • Keep one extraction attempt per invocation and the existing Supervisor backoff. Unavailable providers, embedding timeouts, single-position failures, and hard Worker termination do not silently skip input or receive a misleading success acknowledgement.
  • Cover reopened databases, single-Source recovery, unrelated inference failures, stale attempts, rollback, worker acknowledgement/replay, and exact Source consumption. Update English/Chinese operation guidance, backup restore table lists, RFC text, and real-E2E cleanup.

Are there any user-facing changes?

Repeated Memory generation timeouts reduce subsequent automatic and SDK/HTTP flush windows without operator intervention. memory.window_reduced logs the affected journal range and next limit without Source contents. A single Source may still require changing the provider or timeout; keep the Worker deadline above the provider timeout plus startup/commit overhead.

Schema change is additive: startup creates pc_memory_source_windows, with at most one recovery hint per affected Scope. Existing Source/cursor payload formats and public APIs are unchanged. The new state is an input-size hint, not an acknowledgement of processed evidence or a Supervisor failed-job state. Access uses the Scope primary key rather than scanning pending Sources.

How was this change tested?

  • uv run --locked pytest -q tests/builtin/runtime/test_memory_window_recovery.py tests/builtin/runtime/test_family_processing.py tests/builtin/runtime/test_processing_scheduler.py tests/builtin/persistence/test_cursors.py tests/builtin/persistence/test_mysql_schema.py tests/builtin/persistence/test_processing_migration.py tests/e2e/test_builtin_runtime.py tests/builtin/artifacts/memory/test_capacity.py tests/e2e/test_memory_capacity.py — 75 passed, 13 skipped after rebasing onto the latest Memory-capacity changes. The skipped cases require a live OceanBase database.
  • HTTP/SDK and integration checks: uv run --locked pytest -q tests/test_integration_manifest.py tests/e2e/test_runtime_server.py -k 'memory or source or integration' — 40 passed, 1 live-backend case skipped, 9 deselected on the initial baseline.
  • Negative control: the new reopen/recovery regression fails against unmodified master because the second attempt still selects all four Sources; it passes with this change. SQLite persistence and domain/worker code are real; provider timeouts are injected, not live Qwen measurements.
  • uv lock --locked, Linux-target full ty check, native-target type checking of changed modules/tests, Pydantic AI integration type check, workflow action pin check, and integration-manifest documentation check passed.
  • All non-type pre-commit hooks passed (SKIP=ty-check uv run --locked prek run -a). The default Windows full type check has existing POSIX-only native-code errors, so it is not reported as passing. No live OceanBase/seekdb execution is claimed.

AI usage statement

Implemented and reviewed with OpenAI Codex (GPT-6), including regression tests, concurrency/transaction review, and validation. Execution limitations are stated above.

@Teingi Teingi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Teingi

Teingi commented Sep 30, 2026

Copy link
Copy Markdown
Member

please resolve conflicts

@knqiufan
knqiufan force-pushed the codex/fix-memory-window-timeout-1757 branch from e842fee to a0337e9 Compare October 1, 2026 15:09
@Teingi
Teingi merged commit cd9ba8a into oceanbase:master Oct 1, 2026
23 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: Memory extraction permanently blocked after large Source window timeout — Worker retries identical window without shrinking or advancing Cursor

2 participants