Free DuckDuckGo web search + content extraction — no API key required.
A lightweight Python package that searches the web via DuckDuckGo and extracts readable text from pages. Use it as a CLI tool, a Python library, or a Hermes Agent skill. Powered by the ddgs metasearch library (2.7k ⭐).
- What It Does
- Features
- Installation
- Quick Start
- API Reference
- CLI Reference
- Hermes Agent Integration
- Caching Behavior
- FAQ
- License
web-search performs two main jobs:
-
Web Search — Queries DuckDuckGo via the
ddgsmetasearch library and returns results as{title, url, body}dictionaries. No API key, no account, no credit card — it just works.ddgshandles DuckDuckGo's anti-bot challenges internally so you don't have to. -
Content Extraction — Fetches a URL, strips away all the HTML, JavaScript, navigation menus, and ads, then returns just the readable text. Great for feeding page content into LLMs or doing your own analysis.
You can also combine both in a single call: search for something, then automatically extract the text from every result page. One line of code, done.
| Feature | Details |
|---|---|
| One Dependency | Only needs ddgs (2.7k ⭐) — a battle-tested DuckDuckGo metasearch library that handles anti-bot challenges automatically. |
| No API Key | Uses DuckDuckGo's public interface via ddgs. Free forever. |
| Search Snippets | Returns body (text snippet) for each search result, not just title+URL. |
| Content Extraction | Strips HTML to readable plain text. Skips scripts, styles, nav/footer/header/aside elements. |
| HTML Entity Decoding | Titles come out clean — ' becomes ', & becomes &, etc. |
| Smart Caching | MD5-based file cache in ~/.cache/websearch/. Extract the same URL twice and the second call is instant. |
| Anti-Bot Resilient | DuckDuckGo now challenges direct HTML scrapers with CAPTCHAs. ddgs handles this internally — your searches just work. |
| CLI + Python + Hermes | Three ways to use it, all from the same install. |
| PEP 561 Typed | Ships py.typed marker — compatible with mypy, pylance, and IDE autocompletion. |
pip install git+https://github.com/ChenneyZhuang/web-search.gitThis is the simplest method. It installs the latest version from the main branch.
git clone https://github.com/ChenneyZhuang/web-search.git
cd web-search
pip install -e .The -e flag ("editable install") means changes you make to the source code take effect immediately — great if you plan to modify or contribute.
If you use Hermes Agent, you can install this as a skill:
# First install the package
pip install git+https://github.com/ChenneyZhuang/web-search.git
# Then register the skill
mkdir -p ~/.hermes/skills/web-search
cp SKILL.md ~/.hermes/skills/web-search/Then in Hermes, run /skill web-search to activate it.
- Python 3.9 or newer
ddgs(installed automatically) — the DuckDuckGo metasearch library- Linux or macOS (tested on both; Windows may work but isn't officially tested)
- Internet connection (obviously — it's a web search tool!)
After installation, the websearch command is available in your terminal:
# Basic search — returns titles and URLs
websearch "Canberra weather"
# Specify how many results (default is 5)
websearch "best pizza in Brooklyn" 10
# Search AND extract the text from each page
websearch "DeepSeek API pricing" 3 --extract
# Same thing, short flag
websearch "solar panels Canberra" 5 -eSample output (search-only mode):
1. Canberra Weather Forecast - Bureau of Meteorology
http://www.bom.gov.au/act/forecasts/canberra.shtml
2. Canberra Weather - AccuWeather
https://www.accuweather.com/en/au/canberra/2600/weather-forecast/13285
...
Sample output (with --extract):
============================================================
1. Canberra Weather Forecast - Bureau of Meteorology
http://www.bom.gov.au/act/forecasts/canberra.shtml
============================================================
Canberra area. Partly cloudy. Medium chance of showers, most
likely in the morning and afternoon. Light winds becoming...
from websearch.engine import search, extract, search_and_extract
# ── Search the web ─────────────────────────────────────
results = search("Canberra solar installers", limit=5)
for r in results:
print(f"{r['title']}")
print(f" {r['url']}\n")
# ── Extract text from a single URL ─────────────────────
text = extract("https://en.wikipedia.org/wiki/Canberra")
print(text[:500]) # First 500 characters
# ── Search AND extract in one call ─────────────────────
rich_results = search_and_extract("best cafes Canberra", limit=3)
for r in rich_results:
print(f"## {r['title']}")
print(r['content'][:300])
print("---")All functions live in websearch.engine.
Search DuckDuckGo and return a list of results.
search(query: str, limit: int = 10) -> list[dict]| Parameter | Type | Default | Description |
|---|---|---|---|
query |
str |
(required) | The search query. Can be anything you'd type into DuckDuckGo — natural language, keywords, phrases in quotes, etc. |
limit |
int |
10 |
Maximum number of results to return. DuckDuckGo's HTML page returns around 20–25 results; values above that will just return whatever's available. |
A list of dictionaries, each with three keys:
| Key | Type | Description |
|---|---|---|
title |
str |
The page title as displayed in DuckDuckGo's results. |
url |
str |
The destination URL (direct, no redirect wrapper). |
body |
str |
A text snippet/description from the search result. Useful for quick scanning. |
[
{
"title": "Canberra - Wikipedia",
"url": "https://en.wikipedia.org/wiki/Canberra",
"body": "Canberra is the capital city of Australia..."
},
{
"title": "Visit Canberra - Official Tourism Website",
"url": "https://visitcanberra.com.au/",
"body": "Discover Canberra, Australia's capital city..."
},
...
]- Results are sourced from DuckDuckGo via the
ddgslibrary, which handles anti-bot challenges automatically. No raw HTML scraping anymore — DuckDuckGo now blocks direct scrapers with CAPTCHAs. - The function makes a single API call. No pagination.
- If no results are found, you'll get an empty list
[]. - Network errors from
ddgsare propagated as exceptions (timeout is 15 seconds).
Fetch a URL and extract the readable text content.
extract(url: str, max_chars: int = 5000) -> str| Parameter | Type | Default | Description |
|---|---|---|---|
url |
str |
(required) | The URL of the page to extract text from. Must be a full URL including https://. |
max_chars |
int |
5000 |
Maximum number of characters to return. Text beyond this limit is truncated. Useful for keeping LLM prompts concise. |
A plain-text string with the extracted content. If the page cannot be fetched (network error, 404, etc.), returns a string like [获取失败: <error message>].
The function tries to find the main content area in this order of priority:
<article>tag<main>tag<div>with a class containing "content" or "post"<body>tag (fallback — the whole page)
Once the content area is found, it:
- Removes
<script>,<style>,<nav>,<footer>,<header>, and<aside>elements - Strips all remaining HTML tags
- Collapses whitespace
- Truncates to
max_chars
Results are cached to disk (see Caching Behavior). Calling extract() with the same URL a second time is near-instant.
text = extract("https://example.com/article", max_chars=2000)
print(text)Search the web and extract text from each result — all in one call.
search_and_extract(query: str, limit: int = 5, max_per_page: int = 3000) -> list[dict]| Parameter | Type | Default | Description |
|---|---|---|---|
query |
str |
(required) | The search query. Passed directly to search(). |
limit |
int |
5 |
Maximum number of search results. Passed to search(). |
max_per_page |
int |
3000 |
Maximum characters to extract from each result page. Passed to extract() as max_chars. |
A list of dictionaries, each with three keys:
| Key | Type | Description |
|---|---|---|
title |
str |
The page title. |
url |
str |
The cleaned URL. |
content |
str |
The extracted text content (up to max_per_page characters). |
results = search_and_extract("Python asyncio tutorial", limit=3, max_per_page=2000)
for r in results:
print(r["title"])
print(r["url"])
print(r["content"][:200])
print("---")This function makes 1 + limit HTTP requests (one for the search, then one for each result). With limit=5 and a 15-second timeout per request, worst-case total time is ~90 seconds, though typical pages respond in 1–3 seconds.
websearch QUERY [LIMIT] [--extract | -e]
| Argument | Required | Description |
|---|---|---|
QUERY |
Yes | The search query. Wrap in quotes if it contains spaces. |
LIMIT |
No | Number of results (default: 5). Must be a plain integer. |
--extract, -e |
No | If present, also extract and display text from each result page. |
# Simple search, default 5 results
websearch "capital of Australia"
# 10 results
websearch "how to make sourdough bread" 10
# 3 results with extracted content
websearch "best noise-cancelling headphones 2026" 3 --extract
# Same, using short flag
websearch "top sci-fi books" 5 -e| Code | Meaning |
|---|---|
0 |
Success (results found and displayed). |
1 |
No query provided (help text shown). |
web-search was designed with Hermes Agent in mind. Two ways to use it:
Hermes ships with a built-in ddgs web search backend since v4.17. Just set in ~/.hermes/config.yaml:
web:
backend: ddgsThen restart the gateway. No extra install needed — ddgs is the default free backend replacing Firecrawl. This is the recommended approach: full integration, no skills to load, automatic tool availability.
If you're on an older Hermes version or want content extraction alongside search:
-
Install the Python package:
pip install git+https://github.com/ChenneyZhuang/web-search.git
-
Register the skill file:
mkdir -p ~/.hermes/skills/web-search cp SKILL.md ~/.hermes/skills/web-search/
-
Activate in Hermes chat:
/skill web-search
Once the skill is active, Hermes can call it through its execute_code or terminal tools. The skill tells Hermes to prefer this tool over the built-in search.
Example Hermes interactions:
User: "Search for the latest news about the James Webb Space Telescope"
Hermes (using the skill):
from websearch.engine import search_and_extract results = search_and_extract("James Webb Space Telescope latest news 2026", limit=3)Then summarizes the results for the user.
Via terminal in Hermes:
python3 -c "from websearch.engine import search; import json; print(json.dumps(search('Canberra weather', 5), indent=2))"- Free — no API key, no rate limits, no quota to exhaust
- Redundant — works even when Hermes's search backend is down
- Content extraction — gets the actual page text, not just snippets
- Caching — repeated extractions are instant
extract() uses a simple file-based cache to avoid re-fetching the same URL repeatedly.
- When you call
extract(url), the function computes an MD5 hash of the URL and takes the first 12 hex characters as the cache key. - It checks
~/.cache/websearch/{key}.txtfor a cached result. - If the cache file exists, it returns the cached text immediately (no network request).
- If not, it fetches the page, extracts the text, writes it to the cache file, and returns it.
~/.cache/websearch/
Example:
~/.cache/websearch/a1b2c3d4e5f6.txt
~/.cache/websearch/9f8e7d6c5b4a.txt
-
To clear the cache entirely:
rm -rf ~/.cache/websearch/The directory is recreated automatically on the next
extract()call. -
To clear a specific URL's cache:
import hashlib key = hashlib.md5("https://example.com/page".encode()).hexdigest()[:12] # Then delete ~/.cache/websearch/{key}.txt
-
Cache never expires automatically. If a page changes, you must clear the cache manually for that URL. In practice, this is rarely an issue for short-lived LLM sessions — the freshness of results is usually "good enough" for one conversation.
Only extract() uses the cache. search() always makes a fresh request to DuckDuckGo, so search results are always current.
No. This library uses the free ddgs metasearch library, which queries DuckDuckGo's public interface. No account, no key, no registration required.
DuckDuckGo permits reasonable, low-volume queries. For personal and research use, this is fine. If you plan to make high-volume automated queries, review DuckDuckGo's terms of service and consider using their official Instant Answer API instead.
Because DuckDuckGo now blocks direct HTML scrapers with CAPTCHA challenges. As of mid-2026, hitting html.duckduckgo.com with a plain HTTP request returns "Unfortunately, bots use DuckDuckGo too" instead of search results. The ddgs library handles these anti-bot measures internally so you don't have to.
It hasn't been tested on Windows, but since it uses only standard-library modules (urllib, pathlib, hashlib, re), it should work. The cache directory uses Path.home() / ".cache" / "websearch", which resolves correctly on Windows (C:\Users\<name>\.cache\websearch\). If you try it and run into issues, please open a GitHub issue.
Increase the max_chars parameter:
# Get up to 20,000 characters
text = extract("https://example.com/long-article", max_chars=20000)Or in the CLI, you can't customize max_chars directly — but you can use the Python API for full control.
The function returns a string like [获取失败: HTTP Error 403: Forbidden] instead of raising an exception. This is intentional — when processing multiple pages, one failure shouldn't crash your entire operation. The content will just be the error message.
Yes! Pass your query in any language DuckDuckGo supports:
search("天气 北京", limit=5) # Chinese
search(" météo Paris", limit=5) # French
search("東京 天気", limit=5) # JapaneseThe query is URL-encoded automatically.
The duckduckgo_search package (by deedy5) is a more feature-rich alternative with support for images, videos, news, and instant answers. web-search is intentionally minimal: zero dependencies, ~85 lines of code, and focused on just text search + content extraction. Choose based on your needs.
Each result page is fetched sequentially (not in parallel) with a 15-second timeout. With limit=5, that's up to 5 extra HTTP requests. If speed matters, use search() to get the URLs first, then selectively extract() only the pages you care about.
| Version | Date | Changes |
|---|---|---|
| 0.2.0 | 2026-06-01 | Switch from raw DuckDuckGo HTML scraping to ddgs library. DuckDuckGo now blocks direct scrapers with CAPTCHAs — ddgs handles this. Added body field to search results (text snippets). |
| 0.1.0 | 2026-05-20 | Initial release. Zero-dependency DuckDuckGo HTML search + content extraction. |
MIT — see the LICENSE file for the full text.
Built with ❤️ by Chenney Zhuang.