A small health check for scheduled scrapers and JSON data feeds. It catches silent failures before stale or incomplete data reaches your users.
It checks:
- newest-record age
- empty or undersized datasets
- duplicate IDs
- invalid timestamps
- sudden row-count drops
- previous-run baselines stored automatically
It can save a health record and then fail an unhealthy Apify run, so you can use Apify task failure notifications or webhooks without putting alert credentials inside the Actor.
Run Scraper Freshness Watchdog on Apify
Use examples/input.sample.json to check a two-row feed against a ten-row baseline. The generated examples/output.sample.json reports all four observed issues: too few rows, stale newest timestamp, duplicate IDs, and an 80% row-count drop.
The output file in this repository was generated by the real Actor analyzer with a fixed timestamp. It is deterministic, not a hand-written success claim.
- Create an Apify task for the Actor.
- Supply either inline rows or a public HTTPS JSON URL.
- Set
timestampField,maxAgeMinutes, and your minimum row count. - Enable
autoBaselineto remember row counts between runs. - Enable
failOnIssuesand attach an Apify failure notification or webhook. - Schedule the task at the cadence your data should refresh.
See TUTORIAL.md for a complete setup and response-handling guide.
The current Store listing is pay per platform usage. The Actor performs one bounded health check per run and caps the default run to one output item. Normal Apify platform usage applies. Check the live Store page for the current pricing shown to your account.
It only accepts inline JSON rows or public HTTPS JSON. It rejects local/private-network targets and does not log in, use proxies, or bypass access controls. Use it only with data you are allowed to monitor.
Open a GitHub issue with the source shape, timestamp field, expected update cadence, and the health condition you need. Do not include private URLs, credentials, or customer data.