From fd598d3fd0039f952f023dcbb6ba7df6553ae244 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 14:57:33 -0500 Subject: [PATCH 01/13] fix(search): match MCP highlights and developer category Reject the github search category, document default web/news highlights, and describe developer coverage as public repositories. --- README.md | 51 +++++++++++++-------------- skills/firecrawl-search/SKILL.md | 4 +-- src/__tests__/cli-argv.test.ts | 42 ++++++++++++++++++++++ src/__tests__/commands/search.test.ts | 11 +++--- src/index.ts | 12 ++++--- src/types/search.ts | 9 ++--- 6 files changed, 86 insertions(+), 43 deletions(-) diff --git a/README.md b/README.md index efcc3c8e7e..e1d7437ec8 100644 --- a/README.md +++ b/README.md @@ -300,7 +300,7 @@ firecrawl https://example.com --exclude-tags nav,aside,.ad ### `search` - Search the web -Search the web and optionally scrape content from search results. +Search the web and optionally scrape content from search results. Web and news results include query-relevant highlights by default. ```bash # Basic search @@ -318,15 +318,14 @@ firecrawl search "landscape photography" --sources images # Multiple sources firecrawl search "machine learning" --sources web,news,images -# Filter by category (GitHub, research-affiliated websites, PDFs) -firecrawl search "web data python" --categories github +# Filter by category (research-affiliated websites, PDFs, developer index) firecrawl search "transformer architecture" --categories research -firecrawl search "machine learning" --categories github,research +firecrawl search "machine learning" --categories pdf,research # Note: --categories research narrows *web* results to research-affiliated # websites. To search papers themselves, use `firecrawl research search-papers`. -# Developer search: GitHub issues, merged PRs, READMEs, and docs +# Developer search: public repositories, GitHub issues, merged PRs, READMEs, and docs firecrawl search "axum middleware ordering" --categories developer # Time-based search @@ -347,23 +346,23 @@ firecrawl search "AI data tools" #### Search Options -| Option | Description | -| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `--limit ` | Maximum results (default: 5, max: 100) | -| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | -| `--categories ` | Comma-separated: `github`, `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | -| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | -| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | -| `--country ` | ISO country code (default: US) | -| `--timeout ` | Timeout in milliseconds (default: 60000) | -| `--highlights` | Return query-relevant highlights for each result | -| `--no-highlights` | Keep the original search snippets | -| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | -| `--scrape` | Enable scraping of search results | -| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | -| `--only-main-content` | Include only main content when scraping (default: true) | -| `-o, --output ` | Save to file | -| `--json` | Output as compact JSON | +| Option | Description | +| ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `--limit ` | Maximum results (default: 5, max: 100) | +| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | +| `--categories ` | Comma-separated: `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | +| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | +| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | +| `--country ` | ISO country code (default: US) | +| `--timeout ` | Timeout in milliseconds (default: 60000) | +| `--highlights` | Query-relevant highlights for web and news results (default). Omitted for zero-data-retention searches | +| `--no-highlights` | Keep the original search snippets | +| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | +| `--scrape` | Enable scraping of search results | +| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | +| `--only-main-content` | Include only main content when scraping (default: true) | +| `-o, --output ` | Save to file | +| `--json` | Output as compact JSON | #### Examples @@ -371,8 +370,8 @@ firecrawl search "AI data tools" # Research a topic with recent results firecrawl search "React Server Components" --tbs qdr:m --limit 10 -# Find GitHub repositories -firecrawl search "web data library" --categories github --limit 20 +# Answer a programming question from public repositories, issues, merged PRs, READMEs, and docs +firecrawl search "web data library" --categories developer --limit 20 # Search and get full content firecrawl search "firecrawl documentation" --scrape --scrape-formats markdown --json -o results.json @@ -383,7 +382,7 @@ firecrawl research search-papers "large language models" --json # Narrow web results to research-affiliated websites (not the paper index) firecrawl search "large language models" --categories research --json -# Answer a programming question from issues, merged PRs, READMEs, and docs +# Answer a programming question from public repositories, issues, merged PRs, READMEs, and docs firecrawl search "tokio select cancellation safety" --categories developer --json # Search with location targeting @@ -397,7 +396,7 @@ firecrawl search "AI startups funding" --sources news --tbs qdr:w --limit 15 ### `developer` - Search developer sources -Search an index built for coding agents: GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Use it for a programming question: code behaviour, a library or framework, an API contract, an error message, or a known bug. +Search an index built for coding agents: public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Use it for a programming question: code behaviour, a library or framework, an API contract, an error message, or a known bug. The CLI intentionally keeps this agent-facing surface lean: it accepts only the query and result count. Express repository, source, result-kind, language, topic, license, and other scoping intent in the query text; semantic retrieval handles the scoping. For advanced filters, use the [Developer Index REST API](https://docs.firecrawl.dev/features/developer). diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index 7831635fac..bc354dc425 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -26,7 +26,7 @@ firecrawl search "your query" --sources news --tbs qdr:d -o .firecrawl/news.json Run `firecrawl search --help` for the full option list. -`--categories developer` weighs the developer index beside ordinary web results in this same call (no passage control, no index filters). `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). +`--categories developer` weighs the developer index of public repositories, GitHub issues, merged PRs, READMEs, and docs beside ordinary web results in this same call (no passage control, no index filters). `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). **Done when:** results are saved under `.firecrawl/`, verified non-empty, processed for the request, and one feedback event is sent within the time window (unless opted out). @@ -42,7 +42,7 @@ If no returned tool covers the country/market/segment or required inputs, contin ## Tips -- **`--highlights` on by default:** results are query-relevant excerpts, not full-page snippets. Use `--no-highlights` for the original snippets. +- **`--highlights` on by default:** web and news results are query-relevant excerpts, not full-page snippets. Use `--no-highlights` for the original snippets. Highlights are omitted for zero-data-retention searches. - **`--scrape` fetches full content** — reuse that content instead of re-scraping result URLs. This saves credits and avoids redundant fetches. - Always write results to `.firecrawl/` with `-o` to avoid context window bloat. - Use `jq` to extract URLs or titles: `jq -r '.data.web[].url' .firecrawl/search.json` diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index b3c72d85ad..014a677f59 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -82,11 +82,53 @@ describe('CLI argv parsing', () => { expect(result.stdout).not.toContain(removedFilter); } expect(result.stdout).toContain('scoping intent in'); + expect(result.stdout.replace(/\s+/g, ' ')).toContain('public repositories'); // Lean surface: the CLI does not point at the REST API for filters. expect(result.stdout).not.toContain('docs.firecrawl.dev'); expect(result.stderr).not.toContain('unknown command'); }); + testWithBuiltCli( + 'describes default search highlights and public developer coverage', + () => { + const result = spawnSync( + process.execPath, + [cliPath, 'search', '--help'], + { + cwd: process.cwd(), + encoding: 'utf8', + } + ); + + expect(result.status).toBe(0); + const flattened = result.stdout.replace(/\s+/g, ' '); + expect(flattened).toContain( + 'query-relevant highlights in web and news results by default' + ); + expect(flattened).toContain( + 'Highlights are omitted for zero-data-retention searches' + ); + expect(flattened).toContain('public repositories'); + expect(flattened).toContain('research, pdf, developer'); + expect(flattened).not.toContain('github, research'); + } + ); + + testWithBuiltCli('rejects the legacy github search category', () => { + const result = spawnSync( + process.execPath, + [cliPath, 'search', 'tokio', '--categories', 'github'], + { + cwd: process.cwd(), + encoding: 'utf8', + } + ); + + expect(result.status).toBe(1); + expect(result.stderr).toContain('Invalid category "github"'); + expect(result.stderr).toContain('research, pdf, developer'); + }); + testWithBuiltCli('lists the research command in root help output', () => { const result = spawnSync(process.execPath, [cliPath, '--help'], { cwd: process.cwd(), diff --git a/src/__tests__/commands/search.test.ts b/src/__tests__/commands/search.test.ts index dab7ca2e5e..10b68aaa42 100644 --- a/src/__tests__/commands/search.test.ts +++ b/src/__tests__/commands/search.test.ts @@ -191,14 +191,14 @@ describe('executeSearch', () => { await executeSearch({ query: 'web scraping python', - categories: ['github'], + categories: ['pdf'], }); expect(mockHttpPost).toHaveBeenCalledWith( '/v2/search', expect.objectContaining({ query: 'web scraping python', - categories: [{ type: 'github' }], + categories: [{ type: 'pdf' }], }) ); }); @@ -406,7 +406,7 @@ describe('executeSearch', () => { query: 'comprehensive test', limit: 20, sources: ['web', 'news'], - categories: ['github'], + categories: ['developer'], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -421,7 +421,7 @@ describe('executeSearch', () => { limit: 20, integration: 'cli', sources: [{ type: 'web' }, { type: 'news' }], - categories: [{ type: 'github' }], + categories: [{ type: 'developer' }], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -745,8 +745,7 @@ describe('executeSearch', () => { }); it('should accept valid category types', async () => { - const categoryList: Array<'github' | 'research' | 'pdf' | 'developer'> = [ - 'github', + const categoryList: Array<'research' | 'pdf' | 'developer'> = [ 'research', 'pdf', 'developer', diff --git a/src/index.ts b/src/index.ts index 393197ddce..84bcfd5565 100644 --- a/src/index.ts +++ b/src/index.ts @@ -919,7 +919,9 @@ Max upload size: 50 MB */ function createSearchCommand(): Command { const searchCmd = new Command('search') - .description('Search the web and discover relevant Alexandria tools') + .description( + 'Search the web and discover relevant Alexandria tools, with query-relevant highlights in web and news results by default' + ) .argument('', 'Search query, or alexandria for semantic tool search') .argument('[tool-query]', 'Query for search alexandria') .option( @@ -933,7 +935,7 @@ function createSearchCommand(): Command { ) .option( '--categories ', - 'Comma-separated categories to filter: github, research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches indexed GitHub issues, merged PRs, READMEs, and docs)' + 'Comma-separated categories to filter: research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches an index of public repositories, GitHub issues, merged PRs, READMEs, and docs)' ) .option( '--tbs ', @@ -959,7 +961,7 @@ function createSearchCommand(): Command { ) .option( '--highlights', - 'Return query-relevant highlights for each search result' + 'Return query-relevant highlights for each search result. Highlights are omitted for zero-data-retention searches.' ) .option( '--no-highlights', @@ -1036,7 +1038,7 @@ function createSearchCommand(): Command { .map((c: string) => c.trim().toLowerCase()) as SearchCategory[]; // Validate categories - const validCategories = ['github', 'research', 'pdf', 'developer']; + const validCategories = ['research', 'pdf', 'developer']; for (const category of categories) { if (!validCategories.includes(category)) { console.error( @@ -1103,7 +1105,7 @@ function createSearchCommand(): Command { function createDeveloperCommand(): Command { const developerCmd = new Command('developer') .description( - 'Search an index built for coding agents: GitHub issues, merged PRs, repository READMEs, and curated documentation sites. Express repository, source, language, topic, license, and other scoping intent in the query text; semantic retrieval handles the scoping.' + 'Search an index built for coding agents: public repositories, GitHub issues, merged PRs, repository READMEs, and curated documentation sites. Express repository, source, language, topic, license, and other scoping intent in the query text; semantic retrieval handles the scoping.' ) .argument('', 'Natural-language developer question or search phrase') .option( diff --git a/src/types/search.ts b/src/types/search.ts index 882873ccbe..38bb79a1b4 100644 --- a/src/types/search.ts +++ b/src/types/search.ts @@ -5,7 +5,7 @@ import type { ScrapeFormat } from './scrape'; export type SearchSource = 'web' | 'images' | 'news' | 'alexandria'; -export type SearchCategory = 'github' | 'research' | 'pdf' | 'developer'; +export type SearchCategory = 'research' | 'pdf' | 'developer'; export interface SearchOptions { domainTools?: boolean; @@ -19,7 +19,7 @@ export interface SearchOptions { limit?: number; /** Sources to search: web, images, news, alexandria (CLI default: web,alexandria) */ sources?: SearchSource[]; - /** Categories to filter results: github, research, pdf, developer */ + /** Categories to filter results: research, pdf, developer */ categories?: SearchCategory[]; /** Time-based search parameter (e.g., qdr:h, qdr:d, qdr:w, qdr:m, qdr:y) */ tbs?: string; @@ -100,8 +100,9 @@ export interface NewsSearchResult { } /** - * One hit from the `developer` category. The index covers GitHub issues, - * merged pull requests, repository READMEs, and curated documentation sites. + * One hit from the `developer` category. The index covers public repositories, + * GitHub issues, merged pull requests, repository READMEs, and curated + * documentation sites. * `description` holds the matched passage, which runs to several KB. */ export interface DeveloperSearchResult { From 7a65349059e4afd4ab838772c7631c7e63b48f31 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 15:03:17 -0500 Subject: [PATCH 02/13] fix(search): drop extra highlights and developer qualifiers Use the MCP developer sentence and default-highlights wording. Keep the zero-data-retention note on --highlights only, and switch the leftover github category fixture to developer. --- README.md | 4 ++-- skills/firecrawl-search/SKILL.md | 4 ++-- src/__tests__/cli-argv.test.ts | 6 ++---- src/__tests__/commands/search.test.ts | 4 ++-- src/index.ts | 2 +- 5 files changed, 9 insertions(+), 11 deletions(-) diff --git a/README.md b/README.md index e1d7437ec8..4393964965 100644 --- a/README.md +++ b/README.md @@ -300,7 +300,7 @@ firecrawl https://example.com --exclude-tags nav,aside,.ad ### `search` - Search the web -Search the web and optionally scrape content from search results. Web and news results include query-relevant highlights by default. +Search the web and optionally scrape content from search results, with query-relevant highlights by default. ```bash # Basic search @@ -355,7 +355,7 @@ firecrawl search "AI data tools" | `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | | `--country ` | ISO country code (default: US) | | `--timeout ` | Timeout in milliseconds (default: 60000) | -| `--highlights` | Query-relevant highlights for web and news results (default). Omitted for zero-data-retention searches | +| `--highlights` | Return query-relevant highlights for each result | | `--no-highlights` | Keep the original search snippets | | `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | | `--scrape` | Enable scraping of search results | diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index bc354dc425..4ecfae3272 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -26,7 +26,7 @@ firecrawl search "your query" --sources news --tbs qdr:d -o .firecrawl/news.json Run `firecrawl search --help` for the full option list. -`--categories developer` weighs the developer index of public repositories, GitHub issues, merged PRs, READMEs, and docs beside ordinary web results in this same call (no passage control, no index filters). `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). +For a programming question, add `--categories developer`. It searches an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). **Done when:** results are saved under `.firecrawl/`, verified non-empty, processed for the request, and one feedback event is sent within the time window (unless opted out). @@ -42,7 +42,7 @@ If no returned tool covers the country/market/segment or required inputs, contin ## Tips -- **`--highlights` on by default:** web and news results are query-relevant excerpts, not full-page snippets. Use `--no-highlights` for the original snippets. Highlights are omitted for zero-data-retention searches. +- **`--highlights` on by default:** results are query-relevant excerpts, not full-page snippets. Use `--no-highlights` for the original snippets. - **`--scrape` fetches full content** — reuse that content instead of re-scraping result URLs. This saves credits and avoids redundant fetches. - Always write results to `.firecrawl/` with `-o` to avoid context window bloat. - Use `jq` to extract URLs or titles: `jq -r '.data.web[].url' .firecrawl/search.json` diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index 014a677f59..06818a3395 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -102,11 +102,9 @@ describe('CLI argv parsing', () => { expect(result.status).toBe(0); const flattened = result.stdout.replace(/\s+/g, ' '); + expect(flattened).toContain('query-relevant highlights by default'); expect(flattened).toContain( - 'query-relevant highlights in web and news results by default' - ); - expect(flattened).toContain( - 'Highlights are omitted for zero-data-retention searches' + 'Return query-relevant highlights for each search result' ); expect(flattened).toContain('public repositories'); expect(flattened).toContain('research, pdf, developer'); diff --git a/src/__tests__/commands/search.test.ts b/src/__tests__/commands/search.test.ts index 10b68aaa42..1e31558358 100644 --- a/src/__tests__/commands/search.test.ts +++ b/src/__tests__/commands/search.test.ts @@ -191,14 +191,14 @@ describe('executeSearch', () => { await executeSearch({ query: 'web scraping python', - categories: ['pdf'], + categories: ['developer'], }); expect(mockHttpPost).toHaveBeenCalledWith( '/v2/search', expect.objectContaining({ query: 'web scraping python', - categories: [{ type: 'pdf' }], + categories: [{ type: 'developer' }], }) ); }); diff --git a/src/index.ts b/src/index.ts index 84bcfd5565..a701bde92c 100644 --- a/src/index.ts +++ b/src/index.ts @@ -920,7 +920,7 @@ Max upload size: 50 MB function createSearchCommand(): Command { const searchCmd = new Command('search') .description( - 'Search the web and discover relevant Alexandria tools, with query-relevant highlights in web and news results by default' + 'Search the web and discover relevant Alexandria tools, with query-relevant highlights by default' ) .argument('', 'Search query, or alexandria for semantic tool search') .argument('[tool-query]', 'Query for search alexandria') From 0683c1c5b01dd2759fef8f8fcf2b4795badc2dfe Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 15:10:17 -0500 Subject: [PATCH 03/13] docs(search): lead developer category copy with the flag --- skills/firecrawl-search/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index 4ecfae3272..ad03ed8f86 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -26,7 +26,7 @@ firecrawl search "your query" --sources news --tbs qdr:d -o .firecrawl/news.json Run `firecrawl search --help` for the full option list. -For a programming question, add `--categories developer`. It searches an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). +`--categories developer` searches an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. `--categories research` is a website filter, not the paper index. Dedicated skills: [firecrawl-developer-index](../firecrawl-developer-index/SKILL.md) and [firecrawl-research-index](../firecrawl-research-index/SKILL.md). **Done when:** results are saved under `.firecrawl/`, verified non-empty, processed for the request, and one feedback event is sent within the time window (unless opted out). From 3bef25bdd1ae090c79b93fad220377f81209be3d Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 15:38:34 -0500 Subject: [PATCH 04/13] docs(search): drop ZDR from highlights help The CLI has no enterprise flag, so the retention caveat cannot live on a ZDR option. Highlights help now matches the parameter, and the README intro states the default product. --- README.md | 6 +++--- src/__tests__/cli-argv.test.ts | 1 + src/__tests__/commands/search.test.ts | 4 ++-- src/index.ts | 2 +- 4 files changed, 7 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index 4393964965..6f2fcaf188 100644 --- a/README.md +++ b/README.md @@ -300,7 +300,7 @@ firecrawl https://example.com --exclude-tags nav,aside,.ad ### `search` - Search the web -Search the web and optionally scrape content from search results, with query-relevant highlights by default. +Search the web with query highlights and optionally scrape content from search results. ```bash # Basic search @@ -370,7 +370,7 @@ firecrawl search "AI data tools" # Research a topic with recent results firecrawl search "React Server Components" --tbs qdr:m --limit 10 -# Answer a programming question from public repositories, issues, merged PRs, READMEs, and docs +# Answer a programming question from public repositories, GitHub issues, merged PRs, READMEs, and docs firecrawl search "web data library" --categories developer --limit 20 # Search and get full content @@ -382,7 +382,7 @@ firecrawl research search-papers "large language models" --json # Narrow web results to research-affiliated websites (not the paper index) firecrawl search "large language models" --categories research --json -# Answer a programming question from public repositories, issues, merged PRs, READMEs, and docs +# Answer a programming question from public repositories, GitHub issues, merged PRs, READMEs, and docs firecrawl search "tokio select cancellation safety" --categories developer --json # Search with location targeting diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index 06818a3395..c902f248bf 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -106,6 +106,7 @@ describe('CLI argv parsing', () => { expect(flattened).toContain( 'Return query-relevant highlights for each search result' ); + expect(flattened).not.toContain('zero-data-retention'); expect(flattened).toContain('public repositories'); expect(flattened).toContain('research, pdf, developer'); expect(flattened).not.toContain('github, research'); diff --git a/src/__tests__/commands/search.test.ts b/src/__tests__/commands/search.test.ts index 1e31558358..10b68aaa42 100644 --- a/src/__tests__/commands/search.test.ts +++ b/src/__tests__/commands/search.test.ts @@ -191,14 +191,14 @@ describe('executeSearch', () => { await executeSearch({ query: 'web scraping python', - categories: ['developer'], + categories: ['pdf'], }); expect(mockHttpPost).toHaveBeenCalledWith( '/v2/search', expect.objectContaining({ query: 'web scraping python', - categories: [{ type: 'developer' }], + categories: [{ type: 'pdf' }], }) ); }); diff --git a/src/index.ts b/src/index.ts index a701bde92c..4126f78160 100644 --- a/src/index.ts +++ b/src/index.ts @@ -961,7 +961,7 @@ function createSearchCommand(): Command { ) .option( '--highlights', - 'Return query-relevant highlights for each search result. Highlights are omitted for zero-data-retention searches.' + 'Return query-relevant highlights for each search result.' ) .option( '--no-highlights', From 6e344d52b1947568e6a99c7b019f168749b31c1f Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 16:18:08 -0500 Subject: [PATCH 05/13] docs(search): keep Alexandria command copy and simplify highlights Leave Jonathan's search description alone and advertise the highlights default on the flag, without the clunkier query-relevant wording. --- README.md | 2 +- skills/firecrawl-search/SKILL.md | 2 +- src/__tests__/cli-argv.test.ts | 9 +++++++-- src/index.ts | 6 ++---- src/types/search.ts | 2 +- 5 files changed, 12 insertions(+), 9 deletions(-) diff --git a/README.md b/README.md index 6f2fcaf188..2cd6f26c7d 100644 --- a/README.md +++ b/README.md @@ -355,7 +355,7 @@ firecrawl search "AI data tools" | `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | | `--country ` | ISO country code (default: US) | | `--timeout ` | Timeout in milliseconds (default: 60000) | -| `--highlights` | Return query-relevant highlights for each result | +| `--highlights` | Return highlights for each result (default) | | `--no-highlights` | Keep the original search snippets | | `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | | `--scrape` | Enable scraping of search results | diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index ad03ed8f86..48317d67c3 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -42,7 +42,7 @@ If no returned tool covers the country/market/segment or required inputs, contin ## Tips -- **`--highlights` on by default:** results are query-relevant excerpts, not full-page snippets. Use `--no-highlights` for the original snippets. +- **`--highlights` on by default:** results are excerpts from the page, not the original snippets. Use `--no-highlights` for the original snippets. - **`--scrape` fetches full content** — reuse that content instead of re-scraping result URLs. This saves credits and avoids redundant fetches. - Always write results to `.firecrawl/` with `-o` to avoid context window bloat. - Use `jq` to extract URLs or titles: `jq -r '.data.web[].url' .firecrawl/search.json` diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index c902f248bf..9f8afe4f77 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -102,9 +102,14 @@ describe('CLI argv parsing', () => { expect(result.status).toBe(0); const flattened = result.stdout.replace(/\s+/g, ' '); - expect(flattened).toContain('query-relevant highlights by default'); expect(flattened).toContain( - 'Return query-relevant highlights for each search result' + 'Search the web and discover relevant Alexandria tools' + ); + expect(flattened).not.toContain( + 'with query-relevant highlights by default' + ); + expect(flattened).toContain( + 'Return highlights for each search result (default).' ); expect(flattened).not.toContain('zero-data-retention'); expect(flattened).toContain('public repositories'); diff --git a/src/index.ts b/src/index.ts index 4126f78160..b5f3b7ebd2 100644 --- a/src/index.ts +++ b/src/index.ts @@ -919,9 +919,7 @@ Max upload size: 50 MB */ function createSearchCommand(): Command { const searchCmd = new Command('search') - .description( - 'Search the web and discover relevant Alexandria tools, with query-relevant highlights by default' - ) + .description('Search the web and discover relevant Alexandria tools') .argument('', 'Search query, or alexandria for semantic tool search') .argument('[tool-query]', 'Query for search alexandria') .option( @@ -961,7 +959,7 @@ function createSearchCommand(): Command { ) .option( '--highlights', - 'Return query-relevant highlights for each search result.' + 'Return highlights for each search result (default).' ) .option( '--no-highlights', diff --git a/src/types/search.ts b/src/types/search.ts index 38bb79a1b4..e6167dc8ee 100644 --- a/src/types/search.ts +++ b/src/types/search.ts @@ -31,7 +31,7 @@ export interface SearchOptions { timeout?: number; /** Exclude URLs invalid for other Firecrawl endpoints */ ignoreInvalidUrls?: boolean; - /** Return query-relevant highlights instead of the original search snippets */ + /** Return highlights instead of the original search snippets */ highlights?: boolean; /** Output file path */ output?: string; From 647507204f13b1809c696363aa2478dd36593a32 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sat, 19 Sep 2026 16:23:22 -0500 Subject: [PATCH 06/13] docs(search): call default highlights key excerpts --- skills/firecrawl-search/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index 48317d67c3..f39919607c 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -42,7 +42,7 @@ If no returned tool covers the country/market/segment or required inputs, contin ## Tips -- **`--highlights` on by default:** results are excerpts from the page, not the original snippets. Use `--no-highlights` for the original snippets. +- **`--highlights` on by default:** results are key excerpts from the page. Use `--no-highlights` for the original snippets. - **`--scrape` fetches full content** — reuse that content instead of re-scraping result URLs. This saves credits and avoids redundant fetches. - Always write results to `.firecrawl/` with `-o` to avoid context window bloat. - Use `jq` to extract URLs or titles: `jq -r '.data.web[].url' .firecrawl/search.json` From 3fbdf558e3705ae510b8c582071bbf927f963b78 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sun, 20 Sep 2026 22:38:20 -0500 Subject: [PATCH 07/13] fix(search): clarify highlights and preserve github category --- README.md | 45 ++++++++++++++------------- src/__tests__/cli-argv.test.ts | 24 ++------------ src/__tests__/commands/search.test.ts | 11 ++++--- src/index.ts | 6 ++-- src/types/search.ts | 6 ++-- 5 files changed, 37 insertions(+), 55 deletions(-) diff --git a/README.md b/README.md index 8d45999154..619713467d 100644 --- a/README.md +++ b/README.md @@ -281,7 +281,7 @@ firecrawl https://example.com --exclude-tags nav,aside,.ad ### `search` - Search the web -Search the web with query highlights and optionally scrape content from search results. +Search the web with query-relevant highlights by default and optionally scrape content from search results. ```bash # Basic search @@ -299,9 +299,10 @@ firecrawl search "landscape photography" --sources images # Multiple sources firecrawl search "machine learning" --sources web,news,images -# Filter by category (research-affiliated websites, PDFs, developer index) +# Filter by category (GitHub, research-affiliated websites, PDFs) +firecrawl search "web data python" --categories github firecrawl search "transformer architecture" --categories research -firecrawl search "machine learning" --categories pdf,research +firecrawl search "machine learning" --categories github,research # Note: --categories research narrows *web* results to research-affiliated # websites. To search papers themselves, use `firecrawl research search-papers`. @@ -327,23 +328,23 @@ firecrawl search "AI data tools" #### Search Options -| Option | Description | -| ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `--limit ` | Maximum results (default: 5, max: 100) | -| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | -| `--categories ` | Comma-separated: `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | -| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | -| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | -| `--country ` | ISO country code (default: US) | -| `--timeout ` | Timeout in milliseconds (default: 60000) | -| `--highlights` | Return highlights for each result (default) | -| `--no-highlights` | Keep the original search snippets | -| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | -| `--scrape` | Enable scraping of search results | -| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | -| `--only-main-content` | Include only main content when scraping (default: true) | -| `-o, --output ` | Save to file | -| `--json` | Output as compact JSON | +| Option | Description | +| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `--limit ` | Maximum results (default: 5, max: 100) | +| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | +| `--categories ` | Comma-separated: `github`, `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | +| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | +| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | +| `--country ` | ISO country code (default: US) | +| `--timeout ` | Timeout in milliseconds (default: 60000) | +| `--highlights` | Query-relevant highlights for web and news when available (default) | +| `--no-highlights` | Keep the original search snippets | +| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | +| `--scrape` | Enable scraping of search results | +| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | +| `--only-main-content` | Include only main content when scraping (default: true) | +| `-o, --output ` | Save to file | +| `--json` | Output as compact JSON | #### Examples @@ -351,8 +352,8 @@ firecrawl search "AI data tools" # Research a topic with recent results firecrawl search "React Server Components" --tbs qdr:m --limit 10 -# Answer a programming question from public repositories, GitHub issues, merged PRs, READMEs, and docs -firecrawl search "web data library" --categories developer --limit 20 +# Find GitHub repositories +firecrawl search "web data library" --categories github --limit 20 # Search and get full content firecrawl search "firecrawl documentation" --scrape --scrape-formats markdown --json -o results.json diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index 9f8afe4f77..fa7f802769 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -105,34 +105,14 @@ describe('CLI argv parsing', () => { expect(flattened).toContain( 'Search the web and discover relevant Alexandria tools' ); - expect(flattened).not.toContain( - 'with query-relevant highlights by default' - ); expect(flattened).toContain( - 'Return highlights for each search result (default).' + 'Return query-relevant highlights for web and news results when available (default).' ); - expect(flattened).not.toContain('zero-data-retention'); expect(flattened).toContain('public repositories'); - expect(flattened).toContain('research, pdf, developer'); - expect(flattened).not.toContain('github, research'); + expect(flattened).toContain('github, research, pdf, developer'); } ); - testWithBuiltCli('rejects the legacy github search category', () => { - const result = spawnSync( - process.execPath, - [cliPath, 'search', 'tokio', '--categories', 'github'], - { - cwd: process.cwd(), - encoding: 'utf8', - } - ); - - expect(result.status).toBe(1); - expect(result.stderr).toContain('Invalid category "github"'); - expect(result.stderr).toContain('research, pdf, developer'); - }); - testWithBuiltCli('lists the research command in root help output', () => { const result = spawnSync(process.execPath, [cliPath, '--help'], { cwd: process.cwd(), diff --git a/src/__tests__/commands/search.test.ts b/src/__tests__/commands/search.test.ts index 10b68aaa42..dab7ca2e5e 100644 --- a/src/__tests__/commands/search.test.ts +++ b/src/__tests__/commands/search.test.ts @@ -191,14 +191,14 @@ describe('executeSearch', () => { await executeSearch({ query: 'web scraping python', - categories: ['pdf'], + categories: ['github'], }); expect(mockHttpPost).toHaveBeenCalledWith( '/v2/search', expect.objectContaining({ query: 'web scraping python', - categories: [{ type: 'pdf' }], + categories: [{ type: 'github' }], }) ); }); @@ -406,7 +406,7 @@ describe('executeSearch', () => { query: 'comprehensive test', limit: 20, sources: ['web', 'news'], - categories: ['developer'], + categories: ['github'], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -421,7 +421,7 @@ describe('executeSearch', () => { limit: 20, integration: 'cli', sources: [{ type: 'web' }, { type: 'news' }], - categories: [{ type: 'developer' }], + categories: [{ type: 'github' }], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -745,7 +745,8 @@ describe('executeSearch', () => { }); it('should accept valid category types', async () => { - const categoryList: Array<'research' | 'pdf' | 'developer'> = [ + const categoryList: Array<'github' | 'research' | 'pdf' | 'developer'> = [ + 'github', 'research', 'pdf', 'developer', diff --git a/src/index.ts b/src/index.ts index b5f3b7ebd2..5b53041066 100644 --- a/src/index.ts +++ b/src/index.ts @@ -933,7 +933,7 @@ function createSearchCommand(): Command { ) .option( '--categories ', - 'Comma-separated categories to filter: research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches an index of public repositories, GitHub issues, merged PRs, READMEs, and docs)' + 'Comma-separated categories to filter: github, research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches an index of public repositories, GitHub issues, merged PRs, READMEs, and docs)' ) .option( '--tbs ', @@ -959,7 +959,7 @@ function createSearchCommand(): Command { ) .option( '--highlights', - 'Return highlights for each search result (default).' + 'Return query-relevant highlights for web and news results when available (default).' ) .option( '--no-highlights', @@ -1036,7 +1036,7 @@ function createSearchCommand(): Command { .map((c: string) => c.trim().toLowerCase()) as SearchCategory[]; // Validate categories - const validCategories = ['research', 'pdf', 'developer']; + const validCategories = ['github', 'research', 'pdf', 'developer']; for (const category of categories) { if (!validCategories.includes(category)) { console.error( diff --git a/src/types/search.ts b/src/types/search.ts index e6167dc8ee..ad7fd4a756 100644 --- a/src/types/search.ts +++ b/src/types/search.ts @@ -5,7 +5,7 @@ import type { ScrapeFormat } from './scrape'; export type SearchSource = 'web' | 'images' | 'news' | 'alexandria'; -export type SearchCategory = 'research' | 'pdf' | 'developer'; +export type SearchCategory = 'github' | 'research' | 'pdf' | 'developer'; export interface SearchOptions { domainTools?: boolean; @@ -19,7 +19,7 @@ export interface SearchOptions { limit?: number; /** Sources to search: web, images, news, alexandria (CLI default: web,alexandria) */ sources?: SearchSource[]; - /** Categories to filter results: research, pdf, developer */ + /** Categories to filter results: github, research, pdf, developer */ categories?: SearchCategory[]; /** Time-based search parameter (e.g., qdr:h, qdr:d, qdr:w, qdr:m, qdr:y) */ tbs?: string; @@ -31,7 +31,7 @@ export interface SearchOptions { timeout?: number; /** Exclude URLs invalid for other Firecrawl endpoints */ ignoreInvalidUrls?: boolean; - /** Return highlights instead of the original search snippets */ + /** Return query-relevant highlights instead of the original search snippets */ highlights?: boolean; /** Output file path */ output?: string; From 84f3019e9fe1d17dd1a8308947acf5b242532abe Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 09:48:27 -0500 Subject: [PATCH 08/13] fix(search): scope developer skill and remove github category --- README.md | 45 +++++++++++------------ skills/firecrawl-developer-index/SKILL.md | 2 +- src/__tests__/cli-argv.test.ts | 18 ++++++++- src/__tests__/commands/search.test.ts | 11 +++--- src/index.ts | 4 +- src/types/search.ts | 4 +- 6 files changed, 49 insertions(+), 35 deletions(-) diff --git a/README.md b/README.md index 619713467d..62410a8635 100644 --- a/README.md +++ b/README.md @@ -281,7 +281,7 @@ firecrawl https://example.com --exclude-tags nav,aside,.ad ### `search` - Search the web -Search the web with query-relevant highlights by default and optionally scrape content from search results. +Search the web with query-relevant highlights and optionally scrape content from search results. ```bash # Basic search @@ -299,10 +299,9 @@ firecrawl search "landscape photography" --sources images # Multiple sources firecrawl search "machine learning" --sources web,news,images -# Filter by category (GitHub, research-affiliated websites, PDFs) -firecrawl search "web data python" --categories github +# Filter by category (research-affiliated websites, PDFs, developer index) firecrawl search "transformer architecture" --categories research -firecrawl search "machine learning" --categories github,research +firecrawl search "machine learning" --categories pdf,research # Note: --categories research narrows *web* results to research-affiliated # websites. To search papers themselves, use `firecrawl research search-papers`. @@ -328,23 +327,23 @@ firecrawl search "AI data tools" #### Search Options -| Option | Description | -| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `--limit ` | Maximum results (default: 5, max: 100) | -| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | -| `--categories ` | Comma-separated: `github`, `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | -| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | -| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | -| `--country ` | ISO country code (default: US) | -| `--timeout ` | Timeout in milliseconds (default: 60000) | -| `--highlights` | Query-relevant highlights for web and news when available (default) | -| `--no-highlights` | Keep the original search snippets | -| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | -| `--scrape` | Enable scraping of search results | -| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | -| `--only-main-content` | Include only main content when scraping (default: true) | -| `-o, --output ` | Save to file | -| `--json` | Output as compact JSON | +| Option | Description | +| ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `--limit ` | Maximum results (default: 5, max: 100) | +| `--sources ` | Comma-separated: `web`, `images`, `news` (default: web) | +| `--categories ` | Comma-separated: `research` (research-affiliated websites -- for papers use [`research search-papers`](#research---search-research-papers)), `pdf`, `developer` | +| `--tbs ` | Time filter: `qdr:h` (hour), `qdr:d` (day), `qdr:w` (week), `qdr:m` (month), `qdr:y` (year) | +| `--location ` | Geo-targeting (e.g., "Germany", "San Francisco,California,United States") | +| `--country ` | ISO country code (default: US) | +| `--timeout ` | Timeout in milliseconds (default: 60000) | +| `--highlights` | Query-relevant highlights for web and news when available (default) | +| `--no-highlights` | Keep the original search snippets | +| `--ignore-invalid-urls` | Exclude URLs invalid for other Firecrawl endpoints | +| `--scrape` | Enable scraping of search results | +| `--scrape-formats ` | Scrape formats when `--scrape` enabled (default: markdown) | +| `--only-main-content` | Include only main content when scraping (default: true) | +| `-o, --output ` | Save to file | +| `--json` | Output as compact JSON | #### Examples @@ -352,8 +351,8 @@ firecrawl search "AI data tools" # Research a topic with recent results firecrawl search "React Server Components" --tbs qdr:m --limit 10 -# Find GitHub repositories -firecrawl search "web data library" --categories github --limit 20 +# Search public developer sources +firecrawl search "web data library" --categories developer --limit 20 # Search and get full content firecrawl search "firecrawl documentation" --scrape --scrape-formats markdown --json -o results.json diff --git a/skills/firecrawl-developer-index/SKILL.md b/skills/firecrawl-developer-index/SKILL.md index affd8619c2..3f9a6d73f6 100644 --- a/skills/firecrawl-developer-index/SKILL.md +++ b/skills/firecrawl-developer-index/SKILL.md @@ -1,6 +1,6 @@ --- name: firecrawl-developer-index -description: Search issues, merged pull requests, and READMEs from public code repositories, plus curated documentation. Use when the question is how a library or API behaves, what an error means, or whether a bug was fixed; prefer this over a general web page. +description: Search an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Use when a programming question needs external documentation or upstream evidence, not for inspecting, editing, or debugging local code or private repositories. --- # Firecrawl Developer Index diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index fa7f802769..b51c916733 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -109,10 +109,26 @@ describe('CLI argv parsing', () => { 'Return query-relevant highlights for web and news results when available (default).' ); expect(flattened).toContain('public repositories'); - expect(flattened).toContain('github, research, pdf, developer'); + expect(flattened).toContain('research, pdf, developer'); + expect(flattened).not.toContain('github, research'); } ); + testWithBuiltCli('rejects the github search category', () => { + const result = spawnSync( + process.execPath, + [cliPath, 'search', 'tokio', '--categories', 'github'], + { + cwd: process.cwd(), + encoding: 'utf8', + } + ); + + expect(result.status).toBe(1); + expect(result.stderr).toContain('Invalid category "github"'); + expect(result.stderr).toContain('research, pdf, developer'); + }); + testWithBuiltCli('lists the research command in root help output', () => { const result = spawnSync(process.execPath, [cliPath, '--help'], { cwd: process.cwd(), diff --git a/src/__tests__/commands/search.test.ts b/src/__tests__/commands/search.test.ts index dab7ca2e5e..10b68aaa42 100644 --- a/src/__tests__/commands/search.test.ts +++ b/src/__tests__/commands/search.test.ts @@ -191,14 +191,14 @@ describe('executeSearch', () => { await executeSearch({ query: 'web scraping python', - categories: ['github'], + categories: ['pdf'], }); expect(mockHttpPost).toHaveBeenCalledWith( '/v2/search', expect.objectContaining({ query: 'web scraping python', - categories: [{ type: 'github' }], + categories: [{ type: 'pdf' }], }) ); }); @@ -406,7 +406,7 @@ describe('executeSearch', () => { query: 'comprehensive test', limit: 20, sources: ['web', 'news'], - categories: ['github'], + categories: ['developer'], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -421,7 +421,7 @@ describe('executeSearch', () => { limit: 20, integration: 'cli', sources: [{ type: 'web' }, { type: 'news' }], - categories: [{ type: 'github' }], + categories: [{ type: 'developer' }], tbs: 'qdr:w', location: 'Germany', country: 'DE', @@ -745,8 +745,7 @@ describe('executeSearch', () => { }); it('should accept valid category types', async () => { - const categoryList: Array<'github' | 'research' | 'pdf' | 'developer'> = [ - 'github', + const categoryList: Array<'research' | 'pdf' | 'developer'> = [ 'research', 'pdf', 'developer', diff --git a/src/index.ts b/src/index.ts index 5b53041066..6725b4d3f0 100644 --- a/src/index.ts +++ b/src/index.ts @@ -933,7 +933,7 @@ function createSearchCommand(): Command { ) .option( '--categories ', - 'Comma-separated categories to filter: github, research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches an index of public repositories, GitHub issues, merged PRs, READMEs, and docs)' + 'Comma-separated categories to filter: research, pdf, developer (research filters web results to research-affiliated websites -- it is NOT the paper index; for papers use `firecrawl research search-papers`. developer searches an index of public repositories, GitHub issues, merged PRs, READMEs, and docs)' ) .option( '--tbs ', @@ -1036,7 +1036,7 @@ function createSearchCommand(): Command { .map((c: string) => c.trim().toLowerCase()) as SearchCategory[]; // Validate categories - const validCategories = ['github', 'research', 'pdf', 'developer']; + const validCategories = ['research', 'pdf', 'developer']; for (const category of categories) { if (!validCategories.includes(category)) { console.error( diff --git a/src/types/search.ts b/src/types/search.ts index ad7fd4a756..38bb79a1b4 100644 --- a/src/types/search.ts +++ b/src/types/search.ts @@ -5,7 +5,7 @@ import type { ScrapeFormat } from './scrape'; export type SearchSource = 'web' | 'images' | 'news' | 'alexandria'; -export type SearchCategory = 'github' | 'research' | 'pdf' | 'developer'; +export type SearchCategory = 'research' | 'pdf' | 'developer'; export interface SearchOptions { domainTools?: boolean; @@ -19,7 +19,7 @@ export interface SearchOptions { limit?: number; /** Sources to search: web, images, news, alexandria (CLI default: web,alexandria) */ sources?: SearchSource[]; - /** Categories to filter results: github, research, pdf, developer */ + /** Categories to filter results: research, pdf, developer */ categories?: SearchCategory[]; /** Time-based search parameter (e.g., qdr:h, qdr:d, qdr:w, qdr:m, qdr:y) */ tbs?: string; From a0804935137b9d98d744413e5e8125a7402ab46c Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 10:13:45 -0500 Subject: [PATCH 09/13] docs(skills): simplify developer scope and clarify highlights --- skills/firecrawl-developer-index/SKILL.md | 2 +- skills/firecrawl-search/SKILL.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/firecrawl-developer-index/SKILL.md b/skills/firecrawl-developer-index/SKILL.md index 3f9a6d73f6..eb6444d783 100644 --- a/skills/firecrawl-developer-index/SKILL.md +++ b/skills/firecrawl-developer-index/SKILL.md @@ -1,6 +1,6 @@ --- name: firecrawl-developer-index -description: Search an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Use when a programming question needs external documentation or upstream evidence, not for inspecting, editing, or debugging local code or private repositories. +description: Search an index of public repositories, GitHub issues, merged pull requests, repository READMEs, and curated documentation sites. Use when a programming question needs external documentation or upstream evidence. --- # Firecrawl Developer Index diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index f39919607c..65824a368e 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -42,7 +42,7 @@ If no returned tool covers the country/market/segment or required inputs, contin ## Tips -- **`--highlights` on by default:** results are key excerpts from the page. Use `--no-highlights` for the original snippets. +- **`--highlights` on by default:** results are query-relevant excerpts from the page. Use `--no-highlights` for the original snippets. - **`--scrape` fetches full content** — reuse that content instead of re-scraping result URLs. This saves credits and avoids redundant fetches. - Always write results to `.firecrawl/` with `-o` to avoid context window bloat. - Use `jq` to extract URLs or titles: `jq -r '.data.web[].url' .firecrawl/search.json` From cb9cdfaa5b8470570e325bb99a8bf582625df252 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 10:27:51 -0500 Subject: [PATCH 10/13] fix(search): surface highlights in CLI discovery --- skills/firecrawl-search/SKILL.md | 2 +- src/__tests__/cli-argv.test.ts | 4 ++-- src/index.ts | 6 ++++-- 3 files changed, 7 insertions(+), 5 deletions(-) diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index 65824a368e..ead72e09ed 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-search description: | - Web search with full page content. Use when no URL is known: finding sources, articles, or news. For papers use firecrawl-research-index; for library, API, error, or bug questions use firecrawl-developer-index. + Web search with query-relevant page excerpts and optional full-page content. Use when no URL is known: finding sources, articles, or news. For papers use firecrawl-research-index; for library, API, error, or bug questions use firecrawl-developer-index. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) diff --git a/src/__tests__/cli-argv.test.ts b/src/__tests__/cli-argv.test.ts index b51c916733..139c926bd1 100644 --- a/src/__tests__/cli-argv.test.ts +++ b/src/__tests__/cli-argv.test.ts @@ -103,10 +103,10 @@ describe('CLI argv parsing', () => { expect(result.status).toBe(0); const flattened = result.stdout.replace(/\s+/g, ' '); expect(flattened).toContain( - 'Search the web and discover relevant Alexandria tools' + 'Search the web with query-relevant highlights and discover relevant Alexandria tools' ); expect(flattened).toContain( - 'Return query-relevant highlights for web and news results when available (default).' + 'Return query-relevant page excerpts for web and news results when available (default).' ); expect(flattened).toContain('public repositories'); expect(flattened).toContain('research, pdf, developer'); diff --git a/src/index.ts b/src/index.ts index 6725b4d3f0..a1b3b217f4 100644 --- a/src/index.ts +++ b/src/index.ts @@ -919,7 +919,9 @@ Max upload size: 50 MB */ function createSearchCommand(): Command { const searchCmd = new Command('search') - .description('Search the web and discover relevant Alexandria tools') + .description( + 'Search the web with query-relevant highlights and discover relevant Alexandria tools' + ) .argument('', 'Search query, or alexandria for semantic tool search') .argument('[tool-query]', 'Query for search alexandria') .option( @@ -959,7 +961,7 @@ function createSearchCommand(): Command { ) .option( '--highlights', - 'Return query-relevant highlights for web and news results when available (default).' + 'Return query-relevant page excerpts for web and news results when available (default).' ) .option( '--no-highlights', From 32ddd325f987052afe46f607e8499944a37dd8cb Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 11:00:03 -0500 Subject: [PATCH 11/13] docs(skills): clarify search and scrape handoffs --- skills/firecrawl-scrape/SKILL.md | 2 +- skills/firecrawl-search/SKILL.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/firecrawl-scrape/SKILL.md b/skills/firecrawl-scrape/SKILL.md index 16cfa8a03a..3eb8c2ec09 100644 --- a/skills/firecrawl-scrape/SKILL.md +++ b/skills/firecrawl-scrape/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-scrape description: | - Extract a URL's content as clean markdown, including JS-rendered pages. Use whenever the user provides a URL and wants its content; prefer over WebFetch. + Extract content from a known URL as clean markdown, including JS-rendered pages. Use to read a supplied page or search result. Use firecrawl-search when additional web sources are needed. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index ead72e09ed..9359fbfa9d 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-search description: | - Web search with query-relevant page excerpts and optional full-page content. Use when no URL is known: finding sources, articles, or news. For papers use firecrawl-research-index; for library, API, error, or bug questions use firecrawl-developer-index. + Web search with query-relevant page excerpts and optional full-page content. Use to find sources, articles, news, and current information. If excerpts are insufficient, use firecrawl-scrape to read relevant result URLs. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) From ac1b7422e3ee47da756dbe275fdd971f6aa31ecc Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 16:08:21 -0500 Subject: [PATCH 12/13] docs(skills): keep search and scrape descriptions focused --- skills/firecrawl-scrape/SKILL.md | 2 +- skills/firecrawl-search/SKILL.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/skills/firecrawl-scrape/SKILL.md b/skills/firecrawl-scrape/SKILL.md index 3eb8c2ec09..1d96d95c6c 100644 --- a/skills/firecrawl-scrape/SKILL.md +++ b/skills/firecrawl-scrape/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-scrape description: | - Extract content from a known URL as clean markdown, including JS-rendered pages. Use to read a supplied page or search result. Use firecrawl-search when additional web sources are needed. + Extract content from a known URL as clean markdown, including JS-rendered pages. Use to read a supplied page or search result. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) diff --git a/skills/firecrawl-search/SKILL.md b/skills/firecrawl-search/SKILL.md index 9359fbfa9d..52804eedd8 100644 --- a/skills/firecrawl-search/SKILL.md +++ b/skills/firecrawl-search/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-search description: | - Web search with query-relevant page excerpts and optional full-page content. Use to find sources, articles, news, and current information. If excerpts are insufficient, use firecrawl-scrape to read relevant result URLs. + Web search with query-relevant page excerpts and optional full-page content. Use to find sources, articles, news, and current information. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) From 3b9be80595ade5c8d311cb09dce258d66e691f62 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Mon, 21 Sep 2026 16:11:24 -0500 Subject: [PATCH 13/13] docs(skills): restore scrape description outside search scope --- skills/firecrawl-scrape/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/firecrawl-scrape/SKILL.md b/skills/firecrawl-scrape/SKILL.md index 1d96d95c6c..16cfa8a03a 100644 --- a/skills/firecrawl-scrape/SKILL.md +++ b/skills/firecrawl-scrape/SKILL.md @@ -1,7 +1,7 @@ --- name: firecrawl-scrape description: | - Extract content from a known URL as clean markdown, including JS-rendered pages. Use to read a supplied page or search result. + Extract a URL's content as clean markdown, including JS-rendered pages. Use whenever the user provides a URL and wants its content; prefer over WebFetch. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *)