Skip to content

feat: add provider-level multi-instance failover (M5.1 core) - #41

Merged
rogerdigital merged 2 commits into
mainfrom
feat/m5-external-failover
Sep 19, 2026
Merged

rogerdigital merged 2 commits into
mainfrom
feat/m5-external-failover

Conversation

@rogerdigital

Copy link
Copy Markdown
Owner

What

The provider-level half of M5.1: baseURLs config defines an ordered endpoint pool (first entry primary, non-empty list wins over baseURL, at most 8 entries after duplicate collapse). A network error, timeout, or HTTP 502/503/504/429 raises that endpoint's health penalty and the next attempt — and the first attempt of later searches — goes to the healthiest endpoint. Penalties decay with a 30 s half-life, so the primary is sticky and regains the first slot about a minute after its last failure, or immediately after a successful search.

Design: docs/superpowers/specs/2026-09-19-external-failover-design.md · plan: docs/superpowers/plans/2026-09-19-external-failover.md.

Semantics

  • Failover classes are a closed list (network, timeout, 502/503/504/429). contract failures and other 4xx stay terminal — switching endpoints would silently mask a configuration problem, the same reasoning that keeps empty results honest.
  • Pacing is per endpoint (traffic to instance B never waits on instance A's bucket); same-endpoint retries keep the slot acquired for the first attempt, so single-baseURL behavior is unchanged.
  • The LRU cache is keyed by the whole pool: entries served by a now-dead instance still answer repeats.
  • One totalBudgetMs deadline spans every endpoint's attempts; Retry-After still binds 429 backoff; abort wins over everything.
  • available() requires every named entry to be valid — one malformed entry makes the provider honestly unavailable instead of silently skipping it.

Scope

The CLI/state half of M5.1 (multi-entry setup --url, the external endpoint list in state, per-endpoint doctor/tune reporting) is deferred with its state-schema-3 dependency recorded in the design doc, not silently dropped. Configuring a pool today is a manual profile-config edit, documented in the README.

Verification

  • pnpm verify green locally (typecheck, 901 tests passed / 4 opt-in skips, build, pack, packed-CLI check).
  • 15 new tests (session failover classes, penalty steering, sticky recovery under fake timers, pool-keyed cache, per-endpoint pacing; provider precedence/availability/invalid-entry). All pre-existing single-baseURL tests pass unmodified.
  • Real-instance smoke on the built artifact: pool [dead, live] fails over and returns the same results as [live, dead]; subsequent searches keep steering away from the dead primary.

baseURLs config defines an endpoint pool (first entry primary, non-empty
list wins over baseURL). Network/timeout/502/503/504/429 failures raise
the endpoint's health penalty and fail over to the healthiest endpoint;
penalties decay with a 30 s half-life so the primary is sticky and
recovers the first slot about a minute after its last failure or on its
next success. Pacing is per endpoint, the cache is keyed by the pool, and
the totalBudgetMs deadline spans every endpoint's attempts. Contract
failures and other 4xx never fail over; single-baseURL behavior is
unchanged (existing tests pass unmodified).

The CLI/state half of M5.1 (multi-entry setup --url, external endpoint
list in state, per-endpoint doctor reporting) is deferred with its
state-schema dependency recorded in the design doc.
@rogerdigital
rogerdigital merged commit 3ca1bca into main Sep 19, 2026
6 checks passed
@rogerdigital
rogerdigital deleted the feat/m5-external-failover branch September 19, 2026 06:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant