Skip to content

Fix imported playlists failing every download; add coordinated exponential backoff - #125

Merged
sanylax0 merged 1 commit into
mainfrom
claude/ingest-download-backoff
Jul 21, 2026
Merged

sanylax0 merged 1 commit into
mainfrom
claude/ingest-download-backoff

Conversation

@sanylax2

Copy link
Copy Markdown
Collaborator

The bug

Freshly imported playlists fail downloads for every song — sometimes on YouTube, almost always on Spotify.

Root cause

Not one bug — a retry design that amplified the failure instead of absorbing it.

  1. The retries were the amplifier. An import enqueues every track at once, and each ran its own Retry.run loop (0.7s → 1.4s) with no awareness of the others. When YouTube throttled the burst, all in-flight tracks retried into the same throttle window — extending it — then exhausted their ~2s budget within seconds of each other.

  2. Failure was terminal. process() caught everything and set .failed; resumePreparation explicitly skips .failed. One transient window killed the whole playlist, recoverable only by tapping each row by hand.

  3. The Spotify-specific killer. YouTubeSearch.firstVideoID returns nil both when a real results page has no match and when the page carried no ytInitialData at all — a consent interstitial or bot wall, which is exactly what a 50-track burst provokes. Both collapsed into the non-retryable .noPlayableStream → instant permanent failure. Spotify tracks are the only ones that hit that endpoint (YouTube tracks already carry a video ID) and make double the requests. That's the "almost always Spotify" asymmetry.

  4. Secondary. Expired googlevideo URLs return 403; AudioDownloader mapped that to non-retryable downloadFailed and then retried the same dead URL 3× in 1.2s. A fresh resolve was never attempted.

The fix

IngestBackoff (ContinuityCore, new) Pure backoff math with equal jitter — full jitter can return ~0s, the wrong answer against a rate limiter. Four tuned policies: request / rateLimited / track / sourceCooldown.
IngestThrottle (new) Process-wide cool-down so the app backs off as one client. Any throttle signal arms it; every other in-flight request waits it out. 250 ms minimum spacing paces the initial burst. A stray success clears the streak but never yanks an armed cool-down back (that would oscillate).
Retry Gates on the shared throttle, real exponential backoff, 429s on their own slower curve. Bounded at 4 attempts because it holds an ingestLimiter slot while sleeping.
PreparationQueue.process A transient failure now stays .pending and reschedules on a ~5 min curve plus any remaining shared cool-down. Only non-retryable errors or an exhausted budget reach .failed. Re-fetches the Track by ID after the sleep rather than pinning a @Model for minutes. Adds retryFailedTracks(in:playlist:).
YouTubeSearch.PageOutcome Splits .noResults (permanent) from .unreadable (bot wall → retryable).
AudioDownloader Exponential chunk backoff, 403/410 → new .streamURLExpired which process handles by re-resolving once for a fresh signed URL. Truncated/empty bodies reclassified retryable.

Testing

  • ContinuityCore 187 tests pass, 19 new (IngestBackoffTests ×13, YouTubeSearchTests ×6 including a consent-interstitial fixture). All Accelerate-free, so Linux CI picks them up automatically.
  • App target builds clean for the iOS 27 simulator.
  • Device verification is yours: import a Spotify playlist and watch subsystem:com.continuity.app category:ingest — you should see source cool-down Ns and retrying <track> (attempt N/5) where previously every row went straight to .failed.

🤖 Generated with Claude Code

…ntial backoff

A freshly imported playlist enqueues every track at once, and each ran its own
retry loop with no awareness of the others. When YouTube throttled the burst,
all in-flight tracks retried into the same throttle window — extending it — then
exhausted their ~2s budget (0.7s, 1.4s) within seconds of each other. The
retries were the amplifier, not the cure. Failure was then terminal: process()
marked the track `.failed`, and resumePreparation skips `.failed`, so one
transient window killed the whole playlist until each row was tapped by hand.

Spotify hit a second, worse path. YouTubeSearch.firstVideoID returns nil both
when a real results page has no match and when the page carried no
ytInitialData at all — i.e. a consent interstitial or bot wall, which is exactly
what a 50-track burst provokes. Both collapsed into the non-retryable
`.noPlayableStream`, failing the track permanently. Spotify tracks are the only
ones that hit that endpoint (YouTube tracks already have a video ID) and make
double the requests, hence "almost always Spotify, sometimes YouTube".

- IngestBackoff (ContinuityCore): pure backoff math with equal jitter — full
  jitter can return ~0s, the wrong answer against a rate limiter. Four tuned
  policies (request / rateLimited / track / sourceCooldown).
- IngestThrottle: process-wide cool-down so the app backs off as one client.
  Any throttle signal arms it; every other in-flight request waits it out.
  250ms minimum spacing paces the initial burst. A stray success clears the
  streak but never yanks an armed cool-down back, which would oscillate.
- Retry: gates on the shared throttle, real exponential backoff, 429s on their
  own slower curve. Bounded at 4 attempts since it holds an ingestLimiter slot.
- PreparationQueue.process: a transient failure now stays `.pending` and
  reschedules on a ~5 minute curve plus any remaining shared cool-down instead
  of dying. Only non-retryable errors or an exhausted budget reach `.failed`.
  Re-fetches the Track by ID after the sleep rather than pinning a @model for
  minutes. Adds retryFailedTracks(in:playlist:) for a bulk retry.
- YouTubeSearch.PageOutcome splits .noResults (permanent) from .unreadable
  (bot wall → retryable).
- AudioDownloader: exponential chunk backoff, 403/410 → new .streamURLExpired,
  which process handles by re-resolving once for a fresh signed URL. Truncated
  and empty bodies reclassified retryable.

ContinuityCore: 187 tests pass (19 new). App target builds for the simulator.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@sanylax0
sanylax0 merged commit fd26e87 into main Jul 21, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants