Skip to content

Bot-hostile / heavy sites burn the full capture timeout before failing #22

Description

@Samraaj

Observed on prod 2026-07-06: 8 consecutive collectPage timeout failures across two sites (ninjatransfers.com ×4, websets.exa.ai ×4, incl. a web.archive.org fallback attempt). Each attempt occupied a worker slot for the full timeout window before failing cleanly, and pg-boss then retried (retryLimit 2), tripling the cost.

With WORKER_CONCURRENCY=2 on prod this is less painful than it was under the serial queue, but a determined user retrying a hostile site can still monopolize slots for many minutes producing nothing.

Possible directions (not prescriptive):

  • Early bot-wall detection: the capture layer already has blocked/pollution sanity signals post-capture — probing for interstitial/challenge markers in the first seconds could fail fast before the scroll/settle pipeline runs.
  • Failure memory: after N consecutive timeouts for a URL (cache-keyed), fail immediately for some cooldown window instead of re-burning the timeout (the failed-job history is already in the jobs table).
  • Per-attempt budget shaping: first attempt gets the full window; retries get a reduced one, since a site that timed out at full budget rarely succeeds on an identical retry.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions