Tavily-style web research for AI agents - no API keys, no scraping hacks.
infoseek gives LLM agents the same primitives as paid search services:
multi-engine search, LLM-ready context bundles (ask), clean page
extract, and a built-in prompt-injection guard - keyless, polite,
token-efficient.
pip install git+https://github.com/TruftedBug89/infoseekDesigned so even a small local model can use it: every call is sync, takes a
string, returns a string, and never raises (errors come back as [[...]] notes).
import infoseek
# One-line research, no asyncio, no API key
print(infoseek.research("why is redis faster than postgres", budget=2000))
infoseek.find("rust vs go 2026") # search -> ranked results with urls
infoseek.read("https://example.com/article") # page -> clean text
infoseek.deep("llm quantization tradeoffs") # brief -> multi-angle research
infoseek.help() # usage card, call it to re-learn| you want | call |
|---|---|
| links / sources to cite | find() |
| facts to answer a question | research() |
| the text of one known page | read() |
| a broader, slower, multi-angle brief | deep() |
| call | returns | use when |
|---|---|---|
await infoseek.search(q, n=10) |
list[dict] of results |
you want links + snippets with 30-day recency and community consensus |
await infoseek.ask(q, budget=2500) |
context-bundle string | you want text the model can answer from directly |
await infoseek.last30days(q, days=30) |
curated social & consensus brief | you want people's discussions, real-world consensus, sentiment, or prediction odds |
await infoseek.read_url(url) / extract() |
clean markdown/text | you have a URL to fetch without permission prompts (archive & keyless Jina fallback) |
infoseek.scan(text) (sync) |
verdict: ok / suspect / blocked |
you fetched text yourself and want it screened |
For agents that want zero decisions: await infoseek.run(q) routes by
query shape — bare URL → extract, ask: ... → context bundle,
last30days: ... (or questions like "what are people saying about X") → social brief,
error-message text → fixes-first research, version question → compat research, anything
else → search. infoseek.help() returns a cheat sheet of the whole
surface. ask() auto-reformulates and retries once when first-round results
are poor.
Never a bare empty list. Real agent sessions show models abandoning a
tool that returns [] with no guidance, so empty results are handled in
three steps: (1) auto-retry once with the wide mix, (2) if still empty,
return an actionable hint tailored to the engine prefix (code: → grep.app
only indexes popular repos, try gh: or extract raw files; gh: → check
spelling / try wayback:; swarm: → set SEARXNG_URL; plain → try
engines="wide" or suggest(); …), (3) extract() failures return an
explicit [extract: no content retrieved …] note instead of "".
Helpers: search_many([q1, q2]) (concurrent fan-out, one merged list),
smart_search(q) / search(q, expand=True) (auto-reformulates poor
results), search_error(msg), search_compat(q), changelog(project),
suggest(q), status(), selfcheck().
All output is ASCII-safe and token-lean (compact snippets, dedupe, relevance-sentence extraction, budget caps).
| infoseek | Tavily / Exa / Firecrawl | |
|---|---|---|
| API key | never needed | required (paid) |
| Cost | free | metered |
| Sources | 33 keyless engines + prediction markets + archive ladder | 1 index |
| Prompt-injection guard | built in, µs-fast | not included |
Google and traditional search engines index marketing copy, vendor docs, and SEO spam. AI agents researching tools, libraries, or real-world events need to know what practitioners and real people are experiencing right now.
infoseek.last30days(q, days=30, budget=2500) delivers grounded, time-windowed community consensus and social signal:
- Strict Temporal Windowing: Computes exact verification windows (default: past 30 days) and verifies dates to filter out stale discussions.
- Keyless Social Fan-out: Queries Reddit (via Atom RSS + comment parsing), Hacker News (Algolia API), Polymarket (real-money prediction odds & liquidity), Techmeme, Bluesky, and StockTwits concurrently.
- Engagement-Weighted Ranking: Results are boosted by real social validation—Reddit upvotes, comment density, and prediction volume—so high-signal discussions outrank noise.
- Head-to-Head Comparison Mode: Automatically detects
X vs Yqueries (e.g.uv vs poetry,Claude Code vs Cursor) to output side-by-side consensus, comparative sentiment, and strengths/weaknesses. - Comment & Consensus Extraction: Fetches top comments and practitioner quotes (
u/author (score pts): "...") directly into the brief within strict token budgets.
import asyncio, infoseek
# General community consensus brief
brief = asyncio.run(infoseek.last30days("Claude Code", days=14))
print(brief)
# Head-to-head comparison
vs_brief = asyncio.run(infoseek.last30days("uv vs poetry"))
print(vs_brief)Prefix the query to focus a source; no prefix hits the default mix
(bing + ddg + hn + so + reddit + news). site:<domain> auto-routes.
| prefix | source | prefix | source |
|---|---|---|---|
| (none) | DuckDuckGo + HN + SO + Reddit + News | arxiv: |
arXiv papers |
ddg: |
DuckDuckGo only | openalex: / s2: |
scholarly works (OpenAlex) |
hn: |
Hacker News (Algolia API) | pubmed: / pm: |
biomedical (NCBI) |
reddit: |
Reddit (Atom RSS + discussion comments) | doi: |
DOI / citations (Crossref) |
so: |
Stack Overflow / Exchange | wikidata: / wd: |
structured facts (Q-IDs) |
news: |
Google News RSS | gh: |
GitHub repos |
wiki: |
Wikipedia | code: |
code search (grep.app) |
hf: / huggingface: |
Hugging Face models & quants (GGUF) | pypi: / pip: |
Python packages |
crates: / rust: |
Rust crates | npm: / node: |
JS/TS packages |
yt: / youtube: |
YouTube videos | mdn: / docs: |
MDN Web Docs |
marginalia: |
small / indie web | lobsters: |
lobste.rs |
issues: |
GitHub issues + PRs | prs: |
GitHub pull requests |
releases: |
GitHub release notes | changelog: |
changelog finder |
error: |
error message → fixes | compat: |
version compatibility |
wayback:/wb: |
Wayback Machine snapshots | commoncrawl:/cc: |
Common Crawl index |
swarm: |
SearXNG instance swarm | polymarket: / poly: |
Polymarket prediction odds & volume |
techmeme: |
Techmeme breaking tech news | bluesky: / bsky: |
Bluesky social posts & sentiment |
stocktwits: / twits: |
StockTwits ticker & cashtag streams |
engines="wide" switches search() to the maximum-coverage mix
(ddg + swarm + hn + so + news).
Optional keyed engines activate automatically from env vars (never required):
BRAVE_API_KEY → brave:, SERPER_API_KEY → serper:,
SEARXNG_URL → searxng:. Optional GITHUB_TOKEN raises GitHub rate
limits.
# simple (text in / text out)
$ infoseek find "rust vs go" # ranked results
$ infoseek research "why is redis fast" # context to answer from
$ infoseek deep "llm quantization" # multi-angle brief
$ infoseek read https://example.com/article # one page, clean text
$ infoseek help # usage card
# advanced & structured
$ infoseek search "retrieval augmented generation" --n 10 [--json] [--freshness week]
$ infoseek ask "why is redis faster than postgres" --budget 2000 [--json]
$ infoseek last30days "Claude Code" [--days 30] [--budget 2500] [--json]
$ infoseek extract https://example.com/article [--max-chars 10000]
$ infoseek scan --text "..." | --url https://... # exit 2 if blocked
$ infoseek suggest "python asyn"
$ infoseek status # engine health + last errors
$ infoseek selfcheck # unit + live test batteryextract() climbs an access ladder until it gets text:
- live fetch - official-API fast paths first (GitHub/Wikipedia/HN/Reddit/ PyPI/crates), then a robots-respected direct fetch with bot-wall detection
- Wayback Machine - the closest archived snapshot (raw original via the
id_flag; falls back to recent CDX snapshots). The fetch never touches the origin, so it also works when robots.txt disallows, the site bot-blocks scrapers, or the page/domain no longer exists - Jina Reader (optional) - renders JS-heavy pages; only used when
JINA_API_KEYis set (the keyless tier is Cloudflare-gated from most server IPs)
For search coverage of the bot-walled web, swarm: fans out over
community-hosted SearXNG instances (github.com/searxng/searxng) in
parallel - the instances do the scraping for you. Instance list = curated
seed + daily refresh from the official searx.space list; per-instance health
is cached so walled/dead instances quickly cost nothing. Public instances
increasingly bot-wall datacenter IPs (degrades gracefully, never blocks);
for a guaranteed private swarm run your own instance:
docker run -d -p 8888:8080 searxng/searxng # then set SEARXNG_URL=http://localhost:8888wayback: and commoncrawl: search the archives themselves (snapshot
discovery without touching origin sites).
Retrieved web content is untrusted. infoseek.scan(text) runs layered
regex/structural heuristics (~µs per page, no LLM cost) and returns a verdict:
| level | meaning | handling |
|---|---|---|
ok |
clean | pass through |
suspect |
ambiguous signals | included, flagged as untrusted DATA |
blocked |
clear injection attempt | denied: removed from ask() bundles; extract() returns a denial note |
Detection covers hijack directives, role/framing takeover, prompt
exfiltration, jailbreak phrasing, <system>-style markup, homoglyph and
spaced-letter obfuscation, encoded payloads, and directive density. Policy:
INFOSEEK_GUARD=block (default) | warn | off - keep it on block.
- Python library —
import infoseek(this README). - MCP server —
pip install "infoseek[mcp]", then register stdio commandpython -m infoseek.mcp. Exposesread_url,search,ask,last30days,extract,scan,suggest,status,selfcheck,runas native tools in Antigravity / opencode / Claude Code / Cursor / Windsurf / Continue / Goose. Provides prompt-free URL reading (with markdown formatting and keyless Jina Reader) and 30-day social recency search. - Skill — the repo root is a skill layout (
SKILL.md); point skill-aware harnesses at it (seeAGENTS.mdfor per-harness paths).
| env var | default | purpose |
|---|---|---|
INFOSEEK_CACHE |
platform cache dir | cache location (falls back to temp dir if not writable) |
INFOSEEK_INTERVAL |
1.0 s |
per-host minimum request interval |
INFOSEEK_GUARD |
block |
guard policy: block / warn / off |
INFOSEEK_ENGINE_TIMEOUT |
3.5 s |
per-engine timeout |
SEARXNG_URL |
– | your own SearXNG instance (private swarm, preferred) |
JINA_API_KEY |
– | enables Jina Reader as last-resort page renderer |
Search results cache for 30 min, extractions for 7 days, failures for 90 s. Warm calls return in ~10 ms.
- Legal & polite - official APIs first; HTML parsing only where the site server-renders results (DDG, old.reddit); robots.txt gates direct page fetches; per-host throttling; no CAPTCHA bypass, ever.
- Fast - concurrent engines, per-engine caches, failure markers, n-independent cache keys.
- Cheap - everything is measured in tokens: snippet caps, CTA trimming, relevance-sentence extraction, budget-capped bundles.
- AI-friendly - ASCII output, predictable JSON options, explicit failure notes instead of silent empty results.
pip install -e ".[dev]"
pytest # offline suite (guard battery, routing, ranking, simple API, 90+ tests)
pytest -m live # + live engine probes (network required)
infoseek selfcheck # unit checks + live probes of all engines + smoke runsMIT © 2026 TruftedBug89