Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
54 commits
Select commit Hold shift + click to select a range
33f708a
feat(exchange): expose Firecrawl Exchange v2 shapes over MCP
developersdigest Sep 2, 2026
74fec5b
fix(exchange): accept bare-string search sources and harden discover …
developersdigest Sep 2, 2026
526ec2e
docs(exchange): use provider-relative capability addresses in examples
developersdigest Sep 2, 2026
3868477
Expose Alexandria discovery and shared local credentials in MCP
developersdigest Sep 10, 2026
dac4270
feat: add Alexandria discovery and Find Tools to MCP
developersdigest Sep 11, 2026
a1ca0ff
Merge main into Alexandria MCP
developersdigest Sep 11, 2026
629820a
feat(mcp): align Alexandria search and scrape with domainTools and er…
developersdigest Sep 12, 2026
5a8212e
feat(mcp): rename provider execution key to alexandria
developersdigest Sep 12, 2026
433bb5c
feat(mcp): relay Alexandria provider-terms 403 as a human handoff
developersdigest Sep 12, 2026
715af26
docs(mcp): say Alexandria in the remaining user-visible strings
developersdigest Sep 12, 2026
17716af
feat(mcp): call Find Tools as provider firecrawl capability find-tools
developersdigest Sep 12, 2026
ade1a59
fix(mcp): Alexandria instructions for API-key sessions and wording cl…
developersdigest Sep 12, 2026
472c065
Refine MCP Find Tools progressive catalogue navigation
developersdigest Sep 17, 2026
22dd348
Add explicit provider terms review and acceptance tools
developersdigest Sep 17, 2026
a1ff94e
Use the nested provider agreement catalog in MCP tests
developersdigest Sep 17, 2026
c8edb58
Fix Alexandria retry guidance and address MCP review feedback
developersdigest Sep 17, 2026
4f6e5ec
Enable Alexandria tools in default MCP search
developersdigest Sep 17, 2026
f18a8ba
Remove credit-cost language from Alexandria MCP guidance
developersdigest Sep 17, 2026
1d82247
fix: preserve default web search when provider discovery is unavailable
developersdigest Sep 17, 2026
236dd04
Bound MCP responses and recover invalid environment API keys
developersdigest Sep 18, 2026
de09e3d
Measure MCP context usage and bound oversized errors and framing
developersdigest Sep 18, 2026
a023b33
Align MCP search and scrape with progressive CLI discovery
developersdigest Sep 19, 2026
454dafb
Validate live MCP recovery and isolate smoke test credentials
developersdigest Sep 19, 2026
2415633
Explain Alexandria data coverage and progressive search discovery
developersdigest Sep 19, 2026
4b69d37
Expose discovery detail and selected contract inspection
developersdigest Sep 20, 2026
8d92a35
Fix discovery guidance and split Alexandria integration tests
developersdigest Sep 20, 2026
befc6f2
Trim repetitive Alexandria integration coverage
developersdigest Sep 20, 2026
6def6a0
Make MCP test failures explicit and validate error contracts
developersdigest Sep 20, 2026
e5240d3
Accept compact tool discovery detail
developersdigest Sep 20, 2026
9b8e479
Handle mock failures after response headers
developersdigest Sep 20, 2026
16d8233
Default search tool discovery to compact
developersdigest Sep 20, 2026
c873904
Align MCP contract guidance and workflow version forwarding
developersdigest Sep 20, 2026
2fa5397
Keep Alexandria changes scoped and clarify selective contract expansion
developersdigest Sep 20, 2026
bc6dcd2
Explain preserving request identity before large workflow responses
developersdigest Sep 20, 2026
f2f40ae
Prefer selective contracts and combined retained-result projections
developersdigest Sep 20, 2026
75f2c82
Hand off large retained Alexandria results through remote Bash
developersdigest Sep 20, 2026
3e17241
Remove legacy Exchange discovery interface from MCP
developersdigest Sep 20, 2026
f51ae5d
Reduce Alexandria tests to critical integration paths
developersdigest Sep 20, 2026
8cb231b
Address review feedback on versions, guidance and auth assertions
developersdigest Sep 20, 2026
2fbda53
Disclose provider agreements through scrape recovery
developersdigest Sep 20, 2026
a6f2889
Clarify discovery and defer retained-result guidance
developersdigest Sep 20, 2026
5a777a7
Keep result recovery guidance out of search discovery
developersdigest Sep 20, 2026
6f9f036
Guide disabled provider requests to terms inspection
developersdigest Sep 20, 2026
97985bc
Handle retained result shapes and reported workspace expiry
developersdigest Sep 20, 2026
f1a9f7f
Tidy Alexandria tool surface and guidance
developersdigest Sep 20, 2026
4d8a3e4
Merge main into alexandria-mcp
developersdigest Sep 20, 2026
81a820f
docs: remove internal alias mention from search profile
developersdigest Sep 20, 2026
ca120bc
docs: clarify structured data discovery routing
developersdigest Sep 20, 2026
c76262f
Support Alexandria sessions in the existing feedback tool
developersdigest Sep 21, 2026
83ebb41
Fix smoke assertion for updated Alexandria instructions
developersdigest Sep 21, 2026
1397dd0
Reject Alexandria-only fields on job feedback
developersdigest Sep 21, 2026
4395dba
Merge main into Alexandria MCP and retain credit usage coverage
developersdigest Sep 21, 2026
28e44fe
Restore Alexandria auth validation and payload regression coverage
developersdigest Sep 21, 2026
c9d2f14
Accept missing_capability in Alexandria feedback
nickscamara Sep 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# Changelog

## [3.25.0] - Unreleased

### Added

- Alexandria discovery through `firecrawl_search` with `sources: ["alexandria"]` or `sources: [{ "type": "alexandria" }]`. Both profiles normalize the legacy `exchange` source to `alexandria` and return compact tool suggestions by default in `data.tools` with `creditsUsed` preserved.
- `firecrawl_find_tools` on the full surface provides category browsing, compact provider tools, selected full contracts, URL lookup and `nextTool` navigation.
- `firecrawl_scrape` executes one or up to ten calls using `alexandria: [{ provider, capability, options }]` instead of `url`. Results are returned as `{ success, scrape_id, requestId, data: { alexandria, creditsCost } }`. Errors preserve their code and available charge ID. Keyless sessions receive `Alexandria requires an API key on a team with Alexandria access`.
- Provider terms are read and accepted through nested `firecrawl_scrape` capabilities after a blocked request. Acceptance requires the reviewed version and digest, explicit user authorization and `confirmed: true`.
- Successful large Alexandria results use a retained-result handoff above 20,000 estimated tokens when remote access is verified.

## [3.21.4] - 2026-06-23

### Added
Expand Down
209 changes: 195 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -309,21 +309,23 @@ Use this guide to select the right tool for your task:
- **If you need multi-source research that returns structured data, do not know the URLs, or the answer spans several sites** (an entity plus its fields, a list, a dataset): use **agent**
- **If you want to analyze a whole site or section:** use **crawl** (with limits!)
- **If you need interactive browser automation** (click, type, navigate): use **interact** with a URL for a fresh page, or **scrape** + **interact** when you already scraped the page or need tighter scrape control
- **If you need data from a catalogued provider** (Alexandria): search with `sources: ["alexandria"]`, inspect the selected contract with **find_tools**, and execute it with **scrape** `alexandria`

### Quick Reference Table

| Tool | Best for | Returns |
| ------------ | ---------------------------------------------- | ------------------------------ |
| scrape | Single page content | JSON (preferred) or markdown |
| interact | Interact with a URL or scraped page | Execution result + scrapeId for URL mode |
| map | Discovering URLs on a site | URL[] |
| crawl | Multi-page extraction (with limits) | final crawl status/data after internal polling |
| parse | Files and hosted upload refs | markdown, JSON, or document output |
| search | Web search for info | results[] |
| developer | Programming questions over developer sources | results[] with passages |
| agent | Multi-source research, unknown or many sites | JSON (structured data) |
| monitor | Recurring page checks | monitor/check metadata and diffs |
| research | Paper and GitHub repository research | research results and repo matches |
| Tool | Best for | Returns |
| --------- | ---------------------------------------------- | ------------------------------------------------ |
| scrape | Single page content | JSON (preferred) or markdown |
| interact | Interact with a URL or scraped page | Execution result + scrapeId for URL mode |
| map | Discovering URLs on a site | URL[] |
| crawl | Multi-page extraction (with limits) | final crawl status/data after internal polling |
| parse | Files and hosted upload refs | markdown, JSON, or document output |
| search | Web search for info | results[] |
| find_tools | Alexandria catalogue browsing and URL lookup | providers, tool contracts and nextTool navigation |
| developer | Programming questions over developer sources | results[] with passages |
| agent | Multi-source research, unknown or many sites | JSON (structured data) |
| monitor | Recurring page checks | monitor/check metadata and diffs |
| research | Paper and GitHub repository research | research results and repo matches |

### Format Selection Guide

Expand Down Expand Up @@ -496,6 +498,8 @@ Search the web and optionally extract content from search results.

Set `highlights` to `true` to request query-relevant highlights or `false` to keep the original search snippets. Omit it to use the API's default behavior.

Add `"sources": ["alexandria"]` for semantic tool discovery in `data.tools`, optionally mixed with web/news/images. A query is required. Use `firecrawl_find_tools` on the full MCP surface for contextual lookup and progressive disclosure; see [Alexandria Tools](#15-alexandria-tools).

For scientific papers, see [Research Tools](#12-research-tools-firecrawl_research_): they search paper abstracts and full text, while `categories: ["research"]` here filters ordinary web results to research-affiliated websites.

**Returns:**
Expand Down Expand Up @@ -625,7 +629,6 @@ Starts a crawl job, polls until it reaches a terminal state, and returns the fin
}
```


**Returns:**

- Final crawl status and data after internal polling, including `id`, `status`, `completed`, `total`, `creditsUsed`, `expiresAt`, `next`, and `data`. Use the returned `id` with `firecrawl_check_crawl_status` if you need to re-check the job later.
Expand Down Expand Up @@ -955,7 +958,124 @@ Search an index built for coding agents. The index covers GitHub issues, merged

`firecrawl_search` with `categories: ["developer"]` searches the same index beside the web results. Use this tool instead when you want the matched passages, the `skills` filter, or no web results in the response. The search-only endpoint exposes both tools, and the same choice applies there.

### 15. Credit Usage Tool
### 15. Alexandria Tools

Firecrawl Alexandria is a catalogue of data providers reachable through the Firecrawl API with a Firecrawl API key on a team with Alexandria access. Keyless sessions (hosted or local) get `Alexandria requires an API key on a team with Alexandria access`; Alexandria discovery tools are not listed for hosted keyless sessions.

**Semantic discovery (`firecrawl_search`):**

```json
{
"name": "firecrawl_search",
"arguments": {
"query": "podcast conversations about AI agents",
"sources": ["web", "alexandria"],
"domainTools": true,
"limit": 2
}
}
```

`data.tools` defaults to compact suggestions with provider, capability, and
description. Set `toolDetail: "summary"` for metadata and navigation, or
`toolDetail: "full"` for contracts including inputs, response fields and examples. `domainTools: true` adds contextual matches to
query mentions and result URLs in that same array. Check `warning` for unavailable
discovery. Search requires a query and does not accept catalogue traversal filters.
This discovery works on both the full and search-only MCP surfaces.

**Progressive disclosure (`firecrawl_find_tools`, full MCP surface):**

```json
{
"name": "firecrawl_find_tools",
"arguments": {
"providers": ["particle"],
"capabilities": ["podcasts/episodes/search"],
"expand": ["options", "response", "examples"],
"limit": 2
}
}
```

Start with no arguments for categories, then progressively narrow the catalogue:

| Arguments | Result |
| --- | --- |
| `{}` | Categories and short descriptions |
| `{"categories":["podcasts"]}` | Providers in that category |
| `{"categories":["podcasts"],"providers":["particle"]}` | Compact tool names, descriptions, and prices |
| `{"providers":["particle"],"capabilities":["podcasts/episodes/search"]}` | Complete selected inputs, constraints, response, and examples |

No group hop is required. Explicit `level` supports `categories`, `providers`,
`groups`, or `tools`; explicit `expand` selects contract sections. `expand: []`
keeps results compact even when selecting a capability. For broad full contracts,
explicitly request `level: "tools"` and `expand: ["options", "response", "examples"]`.
URLs provide contextual discovery without fetching the page.

Results are in `data.alexandria[0].data`. Follow an item's `nextTool` by calling
its `name` with its `arguments`; the page's `nextTool` advances pagination.
Existing `next` objects remain Alexandria discovery calls usable through
`firecrawl_scrape`. Discovery costs zero credits and never executes the
provider tools. Read the selected full contract before execution.

These controls require the matching Alexandria API deployment.

**Execute (`firecrawl_scrape` with `alexandria`):** pass `alexandria` instead of `url` (exactly one of the two; requestId and timeout may also be supplied). A single call or an array of up to ten calls is accepted.

```json
{
"name": "firecrawl_scrape",
"arguments": {
"alexandria": [
{
"provider": "fred",
"capability": "series/observations",
"options": { "series_id": "CPIAUCSL" }
}
]
}
}
```

**Returns:** `{ success, scrape_id, requestId, data: { alexandria: [...], creditsCost } }`. Each item is either a result (`provider`, `capability`, `creditsCost`, `data`, `records`, `upstreamStatus`) or an `error` with a `code`; the batch never fails as a whole for a provider error and `data.creditsCost` sums the successful items. Alexandria error bodies (403 without Alexandria access, and 402/409 billing statuses) are relayed in-band with their `code`.

Execution generates one `x-request-id` and returns it on success or failure. Retry
the identical payload with that `requestId`; do not create a new ID after an
uncertain outcome. Available credits are checked by the API.

An Alexandria provider whose terms the team has not accepted fails before anything runs with
HTTP 403 and this body:

```json
{
"success": false,
"code": "THIRD_PARTY_DATA_TERMS_REQUIRED",
"error": "An organization admin must accept the benzinga provider's terms (version 2026-09-12-placeholder) before this request can run. Accept them at https://www.firecrawl.dev/app/alexandria/benzinga",
"requiresAction": {
"type": "accept_terms",
"terms": "benzinga",
"version": "2026-09-12-placeholder",
"url": "https://www.firecrawl.dev/app/alexandria/benzinga"
}
}
```

The tool result relays it as an error with `structuredContent` carrying `code`,
`status: 403`, `requestId`, the `requiresAction` object unchanged, and
`next_actions` (`human_action_required` then `retry_same_request`). Accepting
terms is a legal act. Use the returned `nextTool` call to read the agreement through `firecrawl_scrape`
with `alexandria: [{provider: "firecrawl", capability: "terms/show", options: {provider: "<provider>"}}]`.
Present it to the user and obtain explicit authorization to bind their organization
before calling `firecrawl_scrape` with capability `terms/accept` under provider `firecrawl`.
Its options are `provider`, the exact reviewed `version`
and 64-character lowercase hexadecimal `digest`, and `confirmed: true`. A request
for data is not consent. Authority or eligibility errors may require an organization
admin to use `requiresAction.url`. No automatic acceptance or uncertain retries occur.
Send terms calls separately from execution. These are nested capabilities, not top-level MCP tools.
No credits are charged for the blocked retrieval. After confirmed acceptance, call the same
tool again with the identical payload and `requestId`.

### 16. Credit Usage Tool

The tool requires an authenticated Firecrawl account and is read-only.

Expand Down Expand Up @@ -1030,3 +1150,64 @@ Thanks to MCP.so and Klavis AI for hosting and [@gstarwd](https://github.com/gst
## License

MIT License - see LICENSE file for details

### Structured data and large results

Authenticated search defaults to web results, semantic Alexandria tools, and domain matches. Start with the actual question. Use `firecrawl_find_tools` for direct semantic tool lookup, selected contracts, or progressive browsing: categories → providers → compact tools → selected contract. Execute tools through `firecrawl_scrape`; ordinary URL scraping and search never automatically execute provider tools.

For selected contracts, prefer `expand: ["options", "response"]` and request examples only if the input shape is unclear. Inspect related capabilities together and reuse the returned contracts.

For a potentially large workflow result, supply and preserve a top-level `requestId` before execution. That ID remains available even if the client rejects the response. Regular URL scrapes use the returned scrape ID instead.

For a large retained workflow or regular scrape result, call `firecrawl_scrape` with:

```json
{
"alexandria": {
"provider": "firecrawl",
"capability": "bash",
"options": {
"requestId": "<source-request-or-scrape-id>",
"command": "ls -lh"
}
}
}
```

Read `stdout`, `stderr`, `exitCode`, and `workspaceId` in `data.alexandria[0].data`. Continue with `options: {workspaceId, command}` to inspect `response.json` using `jq`, or `document.md` using bounded text commands for regular scrapes. Keep output selective. Source loading must be a standalone call; its nested source ID differs from the top-level execution request ID. Workspaces expire after five idle minutes, and not every result is retained (including ZDR and API-provider workflow payloads). Search IDs are not supported.

For workflow sources, `response.json` preserves the API envelope: select `.data.alexandria[].data`, then the selected contract’s `response.key` when nonempty. Combine related counts and projections in one Bash command when that shape is known, rather than repeatedly inspecting keys.

Successful Alexandria execution responses above 20,000 estimated tokens (serialized UTF-8 bytes divided by four) return a small handoff after remote Bash confirms the complete batch is accessible. Follow `nextTool` to inspect the retained data. This adds one Bash call and starts a workspace with a five-minute idle TTL. The full payload is preserved. If confirmation fails, the original response stays inline. Search, ordinary URL scrape, error responses and Firecrawl utility calls are unchanged. This delivery budget does not measure the client’s remaining context; no process-local result cache is added. When local filesystem tools are available, saving CLI output and reading selected sections is another option.

The CLI discovery sequence maps to these MCP calls:

| Intent | Tool | Arguments |
| --- | --- | --- |
| Web + semantic + domain tools | `firecrawl_search` | `{"query":"<user question>"}` |
| Semantic tools only | `firecrawl_search` | `{"query":"<user question>","sources":["alexandria"]}` |
| Categories | `firecrawl_find_tools` | `{}` |
| Providers in a category | `firecrawl_find_tools` | `{"categories":["<category-id>"]}` |
| Compact provider tools | `firecrawl_find_tools` | `{"providers":["<provider-id>"]}` |
| Selected contract | `firecrawl_find_tools` | `{"providers":["<provider-id>"],"capabilities":["<capability-id>"]}` |
| Execute | `firecrawl_scrape` | `{"alexandria":{"provider":"<provider-id>","capability":"<capability-id>","options":{"<required-field>":"<value>"}}}` |

Use the full MCP surface for Find Tools and execution; the dedicated search-only surface does not expose them. Reuse a complete contract from search when present rather than making another discovery call. Follow returned `nextTool` navigation only when more results are needed.

### Alexandria session feedback

Use the existing `firecrawl_feedback` tool with `endpoint: "alexandria"`:

```json
{
"endpoint": "alexandria",
"rating": "partial",
"requestedWebsite": {
"url": "https://example.com",
"requestedFunctionality": "Find records and download their attachments"
},
"rationale": "Found summaries but could not retrieve attachments"
}
```

This uses authenticated `POST /v2/feedback`, without a job ID, job-age deadline, or credit refund. Optional `providerFeedback` and `capabilityFeedback` arrays describe coverage gaps and execution issues; the tool schema lists supported issue values. A `new_capability_request` requires `requestedFunctionality`; `missing_capability` (the provider exists but lacks the capability) does not. Existing feedback opt-out and authentication controls apply.
12 changes: 12 additions & 0 deletions docs/search-profile.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,18 @@ no request from this surface can ask the API to fetch third-party page content.
The schema and body construction enforce this directly, and contract tests guard
the behavior. No runtime filter is involved.

## Alexandria source

`sources` entries are source names (`web`, `news`, `images`, `alexandria`) or
`{ type }` objects. Alexandria returns
compact tool suggestions by default in `data.tools`; these are catalogue entries, not executed
provider results. Discovery costs no credits and requires an authenticated team
with Alexandria access; keyless sessions get an explanatory error before any request.

Search requires a query and does not accept catalogue browse mode. The full-surface
catalogue tool (`firecrawl_find_tools`) and execution
with `firecrawl_scrape` using `alexandria` are not part of this six-tool surface.

## OAuth

- **Resource identity.** The surface advertises its own protected resource,
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "firecrawl-mcp",
"version": "3.24.1",
"version": "3.25.0",
"description": "MCP server for Firecrawl — search, scrape, and interact with the web, and search scientific papers. Supports both cloud and self-hosted instances. Features include web search, scraping, page interaction, batch processing, LLM-powered content analysis, and research paper search over biomedical and arXiv literature (PubMed, bioRxiv, medRxiv, arXiv) with citation-graph expansion and full-text reading.",
"type": "module",
"mcpName": "io.github.firecrawl/firecrawl-mcp-server",
Expand Down
Loading
Loading