From efb83927096fdf88a0d0642c856bf88df6f79fc6 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Sun, 27 Sep 2026 14:41:40 -0500 Subject: [PATCH 1/3] feat: per-surface server instructions and read-only hosted scrape - Serve one short instructions string per surface (full, search, keyless) from src/instructions.ts. Each fits in 1,024 characters, routes search and scrape in its first 512, and names only tools the session lists. Hosted /v2/mcp selects them per session through a new fastmcp `instructionsForSession` option (pnpm patch), so API-key and OAuth sessions no longer receive the keyless text. Locally, only stdio without a key or FIRECRAWL_API_URL gets the keyless text; the HTTP transport requires one. - Accept Alexandria provider terms in the dashboard. firecrawl_scrape still reads terms with terms/show but refuses every other terms/* capability, and terms errors link an organization admin to requiresAction.url or the data sources settings page. - Annotate hosted firecrawl_scrape with readOnlyHint: true again. Hosted scrape and search scrapeOptions load a named profile with saveChanges: false; saving browser state goes through firecrawl_interact. Local scrape keeps actions and writable profiles and stays readOnlyHint: false. - Keep firecrawl_agent at readOnlyHint: false: the research agent can click, fill forms, and navigate interactive pages. - Bump to 3.26.0. --- CHANGELOG.md | 3 + README.md | 22 +-- docs/search-profile.md | 8 ++ package.json | 2 +- patches/fastmcp@4.3.2.patch | 72 +++++++++- pnpm-lock.yaml | 6 +- src/alexandria.ts | 2 - src/index.ts | 112 +++++++++++---- src/instructions.ts | 33 +++++ tests/helpers/alexandria-metadata.mjs | 6 +- tests/helpers/exchange-mcp.mjs | 1 + tests/helpers/instructions.mjs | 30 ++++ tests/mcp-alexandria-terms.test.mjs | 31 ++-- tests/mcp-description-budget.test.mjs | 2 +- tests/mcp-instructions.test.mjs | 194 ++++++++++++++++++++++++++ tests/mcp-search-profile.test.mjs | 20 ++- tests/mcp-smoke.test.mjs | 81 +++-------- 17 files changed, 495 insertions(+), 130 deletions(-) create mode 100644 src/instructions.ts create mode 100644 tests/helpers/instructions.mjs create mode 100644 tests/mcp-instructions.test.mjs diff --git a/CHANGELOG.md b/CHANGELOG.md index 620fb0e4..3d03b6f7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,9 @@ ### Changed +- Server instructions are one short string per surface (full, search, keyless). Each stays under 1,024 characters, routes search and scrape within its first 512, and names only tools that session lists. Hosted `/v2/mcp` now selects them per session, so a session with an API key or OAuth token gets the full-surface instructions instead of the keyless ones. Locally, only stdio without an API key, OAuth token, or `FIRECRAWL_API_URL` gets the keyless instructions; a self-hosted `FIRECRAWL_API_URL` and the local HTTP transport, which requires credentials or `FIRECRAWL_API_URL`, get the full-surface instructions. +- An organization admin now accepts Alexandria provider terms in the Firecrawl dashboard. `firecrawl_scrape` refuses `terms/accept` and every other `terms/*` capability except `terms/show`, and terms errors link to `requiresAction.url` or the data sources settings page. Reading terms with `terms/show` is unchanged. +- On the hosted server, `firecrawl_scrape` is annotated `readOnlyHint: true` again. There, `firecrawl_scrape` and `firecrawl_search` `scrapeOptions` load a named `profile` without saving changes to it; save browser state with `firecrawl_interact` and `scrapeOptions.profile`. Local `firecrawl_scrape` keeps browser actions and writable profiles and stays `readOnlyHint: false`. - The search surface (`/v2/mcp-search`) now exposes `firecrawl_find_tools` and `firecrawl_scrape` alongside its six search tools, so agents can execute the Alexandria providers that `firecrawl_search` already returns. Both carry surface-scoped descriptions that name only tools registered on that surface, and Alexandria results there omit the `firecrawl_feedback` pointer. See docs/search-profile.md. ### Fixed diff --git a/README.md b/README.md index d9339057..a03729ac 100644 --- a/README.md +++ b/README.md @@ -37,7 +37,7 @@ A Model Context Protocol (MCP) server that brings [Firecrawl](https://github.com - Use `firecrawl_credit_usage` to check credits left or monthly consumption, optionally broken down by API key. - Consider something else when you need to hold a browser session open across many of your own steps with your own retry and termination logic: each `firecrawl_interact` call runs one `prompt` or `code` turn to completion and returns control — the session can persist across calls via `scrapeId` and ends with `firecrawl_interact_stop`, but you cannot drive it interactively step-by-step from the client side within a single call. -This server lists 26 tools when the full profile registers with default settings (feedback tools included, not running in local-keyless mode). Setting `FIRECRAWL_NO_SEARCH_FEEDBACK=1` and/or `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` removes the corresponding feedback tools and reduces this count, as does local keyless startup. For clients with a tool-slot limit: the hosted keyless endpoint (`https://mcp.firecrawl.dev/v2/mcp`, no API key) exposes only 3 — `firecrawl_scrape`, `firecrawl_search`, `firecrawl_parse` — and the dedicated [search-only endpoint](#search-only-endpoint) (`https://mcp.firecrawl.dev/v2/mcp-search`) exposes a fixed set of 8 tools (search, developer and research search, plus Alexandria catalogue lookup and execution). +This server lists 27 tools when the full profile registers with default settings (feedback tools included, not running in local-keyless mode). Setting `FIRECRAWL_NO_SEARCH_FEEDBACK=1` and/or `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` removes the corresponding feedback tools and reduces this count, as does local keyless startup. For clients with a tool-slot limit: the hosted keyless endpoint (`https://mcp.firecrawl.dev/v2/mcp`, no API key) exposes only 3 (`firecrawl_scrape`, `firecrawl_search`, `firecrawl_parse`), and the dedicated [search-only endpoint](#search-only-endpoint) (`https://mcp.firecrawl.dev/v2/mcp-search`) exposes a fixed set of 8 tools (search, developer and research search, plus Alexandria catalogue lookup and execution). ## Installation @@ -416,6 +416,7 @@ Scrape content from a single URL with advanced options. **Branding format:** Extracts comprehensive brand identity (colors, fonts, typography, spacing, logo, UI components) for design analysis or style replication. **Privacy:** Set `redactPII: true` to return content with personally identifiable information redacted. +**Hosted server:** On the hosted server (`CLOUD_SERVICE=true`) scrape is read-only. It takes no browser `actions`, and a named `profile` loads saved browser state without saving changes to it. To save browser state to a profile, open the page with `firecrawl_interact` (see below). **Returns:** @@ -834,6 +835,7 @@ Interact with a fresh URL or with a page that was already opened by `firecrawl_s - Pass `url` to scrape and open a page for interaction in one MCP call. - Pass `scrapeId` to continue interacting with an existing scraped page. - Pass exactly one of `url` or `scrapeId`, plus either `prompt` or `code`. +- To save browser state (cookies, localStorage) to a named profile, pass `url` with `scrapeOptions: { "profile": { "name": "my-profile", "saveChanges": true } }`. The state is saved when `firecrawl_interact_stop` ends the session. **Usage Example:** @@ -1080,16 +1082,14 @@ HTTP 403 and this body: The tool result relays it as an error with `structuredContent` carrying `code`, `status: 403`, `requestId`, the `requiresAction` object unchanged, and `next_actions` (`human_action_required` then `retry_same_request`). Accepting -terms is a legal act. Use the returned `nextTool` call to read the agreement through `firecrawl_scrape` -with `alexandria: [{provider: "firecrawl", capability: "terms/show", options: {provider: ""}}]`. -Present it to the user and obtain explicit authorization to bind their organization -before calling `firecrawl_scrape` with capability `terms/accept` under provider `firecrawl`. -Its options are `provider`, the exact reviewed `version` -and 64-character lowercase hexadecimal `digest`, and `confirmed: true`. A request -for data is not consent. Authority or eligibility errors may require an organization -admin to use `requiresAction.url`. No automatic acceptance or uncertain retries occur. -Send terms calls separately from execution. These are nested capabilities, not top-level MCP tools. -No credits are charged for the blocked retrieval. After confirmed acceptance, call the same +terms is a legal act, so an organization admin accepts them in the Firecrawl dashboard, not +through MCP. Use the returned `nextTool` call to read the agreement through `firecrawl_scrape` +with `alexandria: [{provider: "firecrawl", capability: "terms/show", options: {provider: ""}}]`, +sent separately from provider execution, and present it to the user. An organization admin then +accepts it at `requiresAction.url`, or at https://www.firecrawl.dev/app/settings?tab=data-sources. +`firecrawl_scrape` refuses every other `terms/*` capability, so it makes no account changes. A request +for data is not consent, and no automatic acceptance or uncertain retries occur. +No credits are charged for the blocked retrieval. After the admin confirms acceptance, call the same tool again with the identical payload and `requestId`. ### 16. Credit Usage Tool diff --git a/docs/search-profile.md b/docs/search-profile.md index b841592e..0cc4e432 100644 --- a/docs/search-profile.md +++ b/docs/search-profile.md @@ -67,6 +67,14 @@ names only tools this surface exposes. `firecrawl_find_tools` is registered the same way. Alexandria results on this surface carry no `feedbackTool` pointer, since `firecrawl_feedback` is not registered here. +`firecrawl_scrape` is read-only here (`readOnlyHint: true`): the surface runs in +hosted safe mode, so it takes no browser `actions`, and a named `profile` loads +saved browser state without saving changes to it. Provider terms can be read +with the nested `terms/show` capability. As on the full surface, an organization +admin accepts them in the dashboard: `firecrawl_scrape` refuses every other +`terms/*` capability, and terms errors link to `requiresAction.url` or +https://www.firecrawl.dev/app/settings?tab=data-sources. + ## Alexandria source `sources` entries are source names (`web`, `news`, `images`, `alexandria`) or diff --git a/package.json b/package.json index 4b8b6e34..c7ddc37f 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "firecrawl-mcp", - "version": "3.25.5", + "version": "3.26.0", "description": "MCP server for Firecrawl — search, scrape, and interact with the web, and search scientific papers. Supports both cloud and self-hosted instances. Features include web search, scraping, page interaction, batch processing, LLM-powered content analysis, and research paper search over biomedical and arXiv literature (PubMed, bioRxiv, medRxiv, arXiv) with citation-graph expansion and full-text reading.", "type": "module", "mcpName": "io.github.firecrawl/firecrawl-mcp-server", diff --git a/patches/fastmcp@4.3.2.patch b/patches/fastmcp@4.3.2.patch index f60b5cc4..10832b02 100644 --- a/patches/fastmcp@4.3.2.patch +++ b/patches/fastmcp@4.3.2.patch @@ -1,8 +1,20 @@ diff --git a/dist/FastMCP.d.cts b/dist/FastMCP.d.cts -index 8eba099cc20f7ae5f70060bebb3871c387cfb2de..c3031da662933c366c7320171db47c246cd3191a 100644 +index 8eba099cc20f7ae5f70060bebb3871c387cfb2de..7f20fb3acd55b27c36f0d540a36ad13214a14275 100644 --- a/dist/FastMCP.d.cts +++ b/dist/FastMCP.d.cts -@@ -605,6 +605,8 @@ type Tool = { + status?: number; + }; + instructions?: string; ++ /** ++ * Per-session instructions, chosen from the session's auth when it is ++ * created. Falls back to `instructions` when it returns undefined. ++ */ ++ instructionsForSession?: (auth: T | undefined) => string | undefined; + /** + * Custom logger instance. If not provided, defaults to console. + * Use this to integrate with your own logging system. +@@ -605,6 +610,8 @@ type Tool boolean; @@ -12,10 +24,22 @@ index 8eba099cc20f7ae5f70060bebb3871c387cfb2de..c3031da662933c366c7320171db47c24 execute: (args: StandardSchemaV1.InferOutput, context: Context) => Promise | string | TextContent | void>; name: string; diff --git a/dist/FastMCP.d.ts b/dist/FastMCP.d.ts -index 810df0c5b6511f34685a0651e6df83acd0239e7b..5c82680c32900c1ca39679e1be25e64c3408aee9 100644 +index 810df0c5b6511f34685a0651e6df83acd0239e7b..bd589ce87ec421a73c40b9d60def26af2dcd957c 100644 --- a/dist/FastMCP.d.ts +++ b/dist/FastMCP.d.ts -@@ -605,6 +605,8 @@ type Tool = { + status?: number; + }; + instructions?: string; ++ /** ++ * Per-session instructions, chosen from the session's auth when it is ++ * created. Falls back to `instructions` when it returns undefined. ++ */ ++ instructionsForSession?: (auth: T | undefined) => string | undefined; + /** + * Custom logger instance. If not provided, defaults to console. + * Use this to integrate with your own logging system. +@@ -605,6 +610,8 @@ type Tool boolean; @@ -25,7 +49,7 @@ index 810df0c5b6511f34685a0651e6df83acd0239e7b..5c82680c32900c1ca39679e1be25e64c execute: (args: StandardSchemaV1.InferOutput, context: Context) => Promise | string | TextContent | void>; name: string; diff --git a/dist/chunk-LWU5CQGW.js b/dist/chunk-LWU5CQGW.js -index 474670585c1fff7d9609d0f900d0743df14a7688..f6091cbe8be4ef30d3eaec90c94a90d5a0802bd0 100644 +index 474670585c1fff7d9609d0f900d0743df14a7688..578399064ef064102c1d856ce593df05b170f22f 100644 --- a/dist/chunk-LWU5CQGW.js +++ b/dist/chunk-LWU5CQGW.js @@ -986,6 +986,9 @@ ${error instanceof Error ? error.stack : JSON.stringify(error)}` @@ -63,8 +87,26 @@ index 474670585c1fff7d9609d0f900d0743df14a7688..f6091cbe8be4ef30d3eaec90c94a90d5 let args = void 0; if (tool.parameters) { const parsed = await tool.parameters["~standard"].validate( +@@ -1557,7 +1569,7 @@ var FastMCP = class extends FastMCPEventEmitter { + } + const session = new FastMCPSession({ + auth, +- instructions: this.#options.instructions, ++ instructions: this.#options.instructionsForSession?.(auth) ?? this.#options.instructions, + logger: this.#logger, + name: this.#options.name, + onToolCall: this.#options.onToolCall, +@@ -1732,7 +1744,7 @@ var FastMCP = class extends FastMCPEventEmitter { + ) : this.#tools; + return new FastMCPSession({ + auth, +- instructions: this.#options.instructions, ++ instructions: this.#options.instructionsForSession?.(auth) ?? this.#options.instructions, + logger: this.#logger, + name: this.#options.name, + onToolCall: this.#options.onToolCall, diff --git a/dist/chunk-UYG7NPM6.cjs b/dist/chunk-UYG7NPM6.cjs -index 3b695bf493c4f54fd970e145cf14fdd8effc9c6d..8883d346f9f32943f823a10bbb72d415f394ca6d 100644 +index 3b695bf493c4f54fd970e145cf14fdd8effc9c6d..1dc64ad0b08ad5bd97000f65e98b23cfbb4dcaab 100644 --- a/dist/chunk-UYG7NPM6.cjs +++ b/dist/chunk-UYG7NPM6.cjs @@ -986,6 +986,9 @@ ${error instanceof Error ? error.stack : JSON.stringify(error)}` @@ -102,3 +144,21 @@ index 3b695bf493c4f54fd970e145cf14fdd8effc9c6d..8883d346f9f32943f823a10bbb72d415 let args = void 0; if (tool.parameters) { const parsed = await tool.parameters["~standard"].validate( +@@ -1557,7 +1569,7 @@ var FastMCP = class extends FastMCPEventEmitter { + } + const session = new FastMCPSession({ + auth, +- instructions: this.#options.instructions, ++ instructions: this.#options.instructionsForSession?.(auth) ?? this.#options.instructions, + logger: this.#logger, + name: this.#options.name, + onToolCall: this.#options.onToolCall, +@@ -1732,7 +1744,7 @@ var FastMCP = class extends FastMCPEventEmitter { + ) : this.#tools; + return new FastMCPSession({ + auth, +- instructions: this.#options.instructions, ++ instructions: this.#options.instructionsForSession?.(auth) ?? this.#options.instructions, + logger: this.#logger, + name: this.#options.name, + onToolCall: this.#options.onToolCall, diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 4eea9e6a..8592ab1e 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -6,7 +6,7 @@ settings: patchedDependencies: fastmcp@4.3.2: - hash: 4ce43b72d62ea76fc9a1cc3a6244a4258d3311a6d369cbe171b5186b3139b489 + hash: d233c08a5a4894781807f0c01f5e17261f0afdf148e09c8346c06de2e827371d path: patches/fastmcp@4.3.2.patch importers: @@ -18,7 +18,7 @@ importers: version: 17.2.2 fastmcp: specifier: 4.3.2 - version: 4.3.2(patch_hash=4ce43b72d62ea76fc9a1cc3a6244a4258d3311a6d369cbe171b5186b3139b489) + version: 4.3.2(patch_hash=d233c08a5a4894781807f0c01f5e17261f0afdf148e09c8346c06de2e827371d) firecrawl: specifier: 4.40.0 version: 4.40.0 @@ -2500,7 +2500,7 @@ snapshots: fast-uri@3.1.3: {} - fastmcp@4.3.2(patch_hash=4ce43b72d62ea76fc9a1cc3a6244a4258d3311a6d369cbe171b5186b3139b489): + fastmcp@4.3.2(patch_hash=d233c08a5a4894781807f0c01f5e17261f0afdf148e09c8346c06de2e827371d): dependencies: '@modelcontextprotocol/sdk': 1.29.0(zod@4.4.3) '@standard-schema/spec': 1.0.0 diff --git a/src/alexandria.ts b/src/alexandria.ts index 76f64325..187f6fe9 100644 --- a/src/alexandria.ts +++ b/src/alexandria.ts @@ -63,8 +63,6 @@ export const findToolsSchema = z export const ALEXANDRIA_CATALOGUE_VERTICALS = 'companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more'; -export const ALEXANDRIA_CATALOGUE_SENTENCE = - "Alexandria is Firecrawl's catalogue of data providers and workflows across " + ALEXANDRIA_CATALOGUE_VERTICALS + '; providers return typed, sourced records through published contracts.'; export const ALEXANDRIA_SOURCES_OPT_OUT = 'A search with sources: ["web"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false.'; diff --git a/src/index.ts b/src/index.ts index fe510130..931b81fc 100644 --- a/src/index.ts +++ b/src/index.ts @@ -44,11 +44,15 @@ import { searchQueryIsValid, ALEXANDRIA_SEARCH_LEAD, ALEXANDRIA_CONTRACT_GUIDANCE, - ALEXANDRIA_CATALOGUE_SENTENCE, ALEXANDRIA_CATALOGUE_VERTICALS, ALEXANDRIA_SOURCES_OPT_OUT, ALEXANDRIA_SEARCH_INSTRUCTIONS, } from './alexandria'; +import { + FULL_INSTRUCTIONS, + KEYLESS_INSTRUCTIONS, + SEARCH_INSTRUCTIONS, +} from './instructions'; import { alexandriaOutput } from './alexandria-output'; import { registerDeveloperTools } from './developer'; import { extractSingleTrustedClientIp } from './keyless-client-ip'; @@ -124,6 +128,8 @@ type ServerProfile = { resourceName: string; /** Server-level instructions surfaced to clients. */ instructions: string; + /** Per-session instructions; falls back to `instructions` when unset or undefined. */ + instructionsForSession?: (session?: SessionData) => string | undefined; /** OAuth protected-resource identifier for this surface. */ resourceUrl: string; /** httpStream endpoint override (defaults to fastmcp's own default). */ @@ -1011,6 +1017,26 @@ type TermsRequiredAction = { type ExchangeErrorContext = { tool: string; requestId?: string; providers?: string[] }; +const DATA_SOURCES_SETTINGS_URL = + 'https://www.firecrawl.dev/app/settings?tab=data-sources'; + +// An organization admin accepts provider terms in the dashboard. Through MCP, +// agents only read them with terms/show, so firecrawl_scrape changes no +// account state. +function isTermsWrite(call: { provider: string; capability: string }): boolean { + const capability = call.capability.trim().toLowerCase(); + return ( + call.provider.trim().toLowerCase() === 'firecrawl' && + capability.startsWith('terms/') && + capability !== 'terms/show' + ); +} + +function termsWriteError(): UserError { + const message = `Provider terms are accepted in the Firecrawl dashboard, not through this connection. Read them with terms/show, then ask an organization admin to accept them at ${DATA_SOURCES_SETTINGS_URL}.`; + return new UserError(message, { code: 'invalid_option', status: 400, message }); +} + function termsRequiredAction(body: unknown): TermsRequiredAction | undefined { const data = body as | { code?: unknown; requiresAction?: unknown } @@ -1044,7 +1070,7 @@ function termsRequiredError( const retryIdentity = requestId ? ` and requestId ${requestId}` : ''; const message = `Alexandria provider terms required. An organization admin must accept the ${action.terms} provider's terms (version ${action.version}) before this request can run.`; return new UserError( - `${message}\n\n1. Read the agreement with firecrawl_scrape using alexandria: {provider: "firecrawl", capability: "terms/show", options: {provider: "${action.terms}"}} and present it to the user. Only after explicit authorization to accept that exact version and digest, use firecrawl_scrape with alexandria: {provider: "firecrawl", capability: "terms/accept", options: {provider: "${action.terms}", version: "", digest: "", confirmed: true}}. Send terms calls separately from provider execution. If authority or eligibility requires a dashboard action, ask an organization admin to visit ${action.url} or https://www.firecrawl.dev/app/settings?tab=data-sources if that page is unavailable. Never infer acceptance from a data request.\n2. After they confirm, call ${context.tool} again with the identical payload${retryIdentity}. If the same terms error persists, stop and ask an organization admin to check access at https://www.firecrawl.dev/app/settings?tab=data-sources.\n\nDo not retry until acceptance is confirmed.`, + `${message}\n\n1. Read the agreement with firecrawl_scrape using alexandria: {provider: "firecrawl", capability: "terms/show", options: {provider: "${action.terms}"}} and present it to the user. Send terms/show separately from provider execution. Terms are accepted in the Firecrawl dashboard, not through this connection: ask an organization admin to accept them at ${action.url}, or ${DATA_SOURCES_SETTINGS_URL} if that page is unavailable. Never infer acceptance from a data request.\n2. After the admin confirms, call ${context.tool} again with the identical payload${retryIdentity}. If the same terms error persists, stop and ask an organization admin to check access at ${DATA_SOURCES_SETTINGS_URL}.\n\nDo not retry until acceptance is confirmed.`, { code: TERMS_REQUIRED_CODE, status: 403, @@ -1128,7 +1154,7 @@ async function relayExchangeError( : undefined; if (disabledProvider) { throw new UserError( - `${message} Read the provider terms and status using nextTool and present them to the user. Never infer acceptance from a data request. Only after explicit authorization for the reviewed version and digest may you call firecrawl_scrape with alexandria:{provider:"firecrawl",capability:"terms/accept",options:{provider:"${disabledProvider}",version:"",digest:"",confirmed:true}}. Send terms calls separately. Acceptance may not restore disabled access; an organization admin may need to review https://www.firecrawl.dev/app/settings?tab=data-sources. Retry the original request only after access is restored.`, + `${message} Read the provider terms and status using nextTool and present them to the user. Never infer acceptance from a data request. Terms are accepted in the Firecrawl dashboard, not through this connection; acceptance may not restore disabled access, and an organization admin can review access at ${DATA_SOURCES_SETTINGS_URL}. Retry the original request only after access is restored.`, { code: typeof data?.code === 'string' ? data.code : 'exchange_error', status: 403, @@ -1213,18 +1239,6 @@ const openAiAppsChallengeToken = normalizeHeader( process.env.OPENAI_APPS_CHALLENGE_TOKEN ); -const FULL_PROFILE_INSTRUCTIONS = - `Firecrawl provides web search, page retrieval, site URL discovery, multi-page collection, structured page data, monitoring, and multi-source research that returns structured data. Match the requested operation to the tool boundary: firecrawl_scrape retrieves one supplied page and can return JSON matching a supplied schema, firecrawl_map enumerates URLs under a site without retrieving their content, and firecrawl_agent runs multi-source research and returns structured data when the URLs are not known or the answer spans several sites (an entity plus its fields, a list, a dataset); its result is read with firecrawl_agent_status. Authenticated firecrawl_search returns web results together with matching Alexandria providers in data.tools. ${ALEXANDRIA_CATALOGUE_SENTENCE} A matching provider can return the same fields across several entities, provenance, exact figures or timestamps, or a large set of records through firecrawl_scrape with its published contract. If web results already answer the question, use them. For the same fields across multiple pages, firecrawl_find_tools offers free provider discovery. Use firecrawl_find_tools to read a contract that was not returned in full or to browse the catalogue by category. ${ALEXANDRIA_SOURCES_OPT_OUT} If no provider fits, continue with web search or firecrawl_agent. For biomedical, life-science, clinical, or arXiv literature, the firecrawl_research_* tools search a paper index of abstracts and full text; firecrawl_search with categories: ["research"] is a website filter over ordinary web results and reaches different sources. For a programming question (code behaviour, a library or framework, an API contract, an error message, or a known bug), firecrawl_developer_search (or firecrawl_search with categories: ["developer"]) searches an index of public repositories, GitHub issues, merged pull requests, READMEs, and code documentation. firecrawl_search with sources: [{type: "alexandria"}] returns compact tool summaries in data.tools; toolDetail: "full" includes contracts, firecrawl_find_tools starts with categories, lists providers, then compact tools, and expands the selected full contract, and firecrawl_scrape with alexandria: [{provider, capability, options}] executes up to ten capabilities and returns their results. Alexandria access needs an API key on a team with it enabled. Provide only the required inputs and account for stated network or external side effects.`; -const KEYLESS_PROFILE_INSTRUCTIONS = `Hosted keyless sessions expose firecrawl_search, firecrawl_scrape, and firecrawl_parse with usage limits. firecrawl_search searches the web. For programming questions, firecrawl_search with categories: ["developer"] searches indexed public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation. For biomedical, life-science, clinical, or arXiv literature, firecrawl_search with categories: ["research"] filters ordinary web results to research-affiliated websites. firecrawl_scrape retrieves one supplied page and can return JSON matching a supplied schema. firecrawl_parse processes supported local files through its two-phase upload flow. An Authorization bearer API key can provide higher usage limits and expose additional tools, subject to plan, deployment, and team policy, including firecrawl_map for site URL discovery, firecrawl_agent and firecrawl_agent_status for multi-source research that returns structured data when the URLs are not known, firecrawl_research_* for paper-index and repository research, and firecrawl_find_tools as the progressive Alexandria catalogue lookup alongside the Alexandria options of firecrawl_search and firecrawl_scrape for catalogued data providers.`; - -// The search surface exposes web/developer/research search plus the two Alexandria -// tools (catalogue lookup and provider execution). Its instructions -// and tool copy describe just those tools and stay neutral about how a client -// uses them. -const SEARCH_PROFILE_INSTRUCTIONS = - ALEXANDRIA_SEARCH_INSTRUCTIONS + - ` Firecrawl provides web, developer, and research search, and executes catalogued Alexandria data providers. Use firecrawl_search to find relevant results across the web and specialized indexes; authenticated searches also return matching Alexandria providers in data.tools. firecrawl_find_tools provides catalogue browsing and provider contracts; firecrawl_scrape with an alexandria body executes a selected capability; firecrawl_scrape with a url retrieves one supplied page. For a programming question, firecrawl_developer_search searches indexed public repositories, GitHub issues, merged pull requests, READMEs, and code documentation and returns the matched passages, and skills: "only" narrows it to agent-skill files; firecrawl_search with categories: ["developer"] reaches the same index beside ordinary web results, returning the hits in the web group rather than as passages and offering no skills filter. For a biomedical, life-science, clinical, or arXiv literature question, the firecrawl_research_* tools search the paper index, while categories: ["research"] on firecrawl_search filters ordinary web results to research-affiliated websites. Use the firecrawl_research_* tools to search academic and research literature, expand from anchor papers via the citation graph, and read full-text passages from a specific paper. Search and discovery tools are read-only and return ranked results. Billing: web, developer and research search are billed per request; Alexandria discovery (firecrawl_search with sources ["alexandria"] alone, and firecrawl_find_tools) is free; firecrawl_scrape is billed, as a page retrieval in url mode or at each executed capability's listed price in alexandria mode.`; - // The exact set of tools the search surface exposes. Registration is filtered // against this set, so anything not listed here can never appear on that // instance's tools/list or be called through it. @@ -1261,12 +1275,23 @@ const SEARCH_SURFACE_VARIANT_TOOLS = new Set([ function makeFullProfile(): ServerProfile { const account = getPrimaryEndpoint() === '/v2/mcp-oauth'; - const hasCredential = Boolean(resolveCredentialFromEnv()); + const hosted = process.env.CLOUD_SERVICE === 'true'; return { id: account ? 'account' : 'full', resourceName: account ? 'Firecrawl MCP Account' : 'Firecrawl MCP', + // Locally, only stdio without credentials or FIRECRAWL_API_URL runs + // keyless; the HTTP transport requires one of them. instructions: - account || hasCredential ? FULL_PROFILE_INSTRUCTIONS : KEYLESS_PROFILE_INSTRUCTIONS, + !account && (hosted || isLocalKeylessStartup()) + ? KEYLESS_INSTRUCTIONS + : FULL_INSTRUCTIONS, + // Hosted /v2/mcp lists tools per session (guardHostedTool), so each + // session's instructions name only the tools it lists. + instructionsForSession: + hosted && !account + ? (session) => + listsOnlyKeylessTools(session) ? KEYLESS_INSTRUCTIONS : FULL_INSTRUCTIONS + : undefined, resourceUrl: account ? (normalizeHeader(process.env.FIRECRAWL_MCP_RESOURCE_URL) ?? DEFAULT_MCP_OAUTH_RESOURCE_URL) @@ -1298,7 +1323,7 @@ function makeSearchProfile({ return { id: 'search', resourceName: 'Firecrawl Search', - instructions: SEARCH_PROFILE_INSTRUCTIONS, + instructions: SEARCH_INSTRUCTIONS, resourceUrl: getSearchMcpResourceUrl(), endpoint: primary ? DEFAULT_MCP_SEARCH_ENDPOINT : getSearchMcpEndpoint(), port: primary @@ -1327,6 +1352,7 @@ function createServer(profile: ServerProfile): FastMCP { name: 'firecrawl-fastmcp', version: packageVersion as `${number}.${number}.${number}`, instructions: profile.instructions, + instructionsForSession: profile.instructionsForSession, logger: new ConsoleLogger(), roots: { enabled: false }, oauth: { @@ -1368,6 +1394,12 @@ function isHostedKeylessSession(session?: SessionData): boolean { ); } +// Hosted keyless sessions and sessions with a rejected credential list only +// KEYLESS_TOOL_NAMES (see guardHostedTool). +function listsOnlyKeylessTools(session?: SessionData): boolean { + return Boolean(session?.credentialError) || isHostedKeylessSession(session); +} + // A stdio client without a cloud credential can use only the keyless tools. // Do this at registration time so unsupported feedback tools are not advertised. function isLocalKeylessStartup(): boolean { @@ -1607,9 +1639,7 @@ function guardHostedTool( // leave MCP clients that stop after tools/list unable to ever surface // the recovery guidance; the full non-keyless schema would over-disclose // to a request carrying an unrecognized or invalid credential. - (session?.credentialError || isHostedKeylessSession(session) - ? keylessTool - : true) && + (listsOnlyKeylessTools(session) ? keylessTool : true) && (canList?.(session) ?? true), beforeValidate: async (args: unknown, session: SessionData) => { const code = session?.credentialError @@ -2005,6 +2035,22 @@ const scrapeParamsSchema = z.object({ .optional(), }); +// In safe mode firecrawl_scrape and firecrawl_search are read-only, so a named +// profile they open loads saved browser state without writing it back. +// firecrawl_interact (url with scrapeOptions.profile) saves profile changes. +const readOnlyProfileSchema = z + .object({ name: z.string() }) + .describe('Loads a saved browser profile without saving changes to it.'); + +function withReadOnlyProfile( + options: Record +): Record { + const profile = options.profile as { name: string } | undefined; + return SAFE_MODE && profile + ? { ...options, profile: { name: profile.name, saveChanges: false } } + : options; +} + // firecrawl_scrape accepts either a page URL or an Exchange batch. The base // schema stays url-required because search, crawl, and monitor reuse it for // nested scrapeOptions, where `alexandria` has no meaning. @@ -2012,6 +2058,7 @@ const ALEXANDRIA_IGNORED_SCRAPE_OPTIONS = new Set(['toolDetail', 'domainTools']) const scrapeToolParamsSchema = scrapeParamsSchema .extend({ + ...(SAFE_MODE ? { profile: readOnlyProfileSchema.optional() } : {}), url: z.string().url().optional(), timeout: z.number().int().positive().optional().describe("Execution timeout in milliseconds."), requestId: z @@ -2453,7 +2500,13 @@ const scrapeTool: RegisteredTool = { name: 'firecrawl_scrape', annotations: { title: 'Firecrawl scrape', - readOnlyHint: false, // Alexandria capabilities can record provider agreement acceptance. + // Safe mode (hosted) offers no browser actions and loads named profiles + // without saving changes. Terms acceptance is refused on every surface + // and happens in the dashboard. Alexandria capabilities retrieve provider data, + // and retained-result Bash processes the caller's own result in a + // workspace that expires after five idle minutes. None of these change + // account, provider, or website state. + readOnlyHint: SAFE_MODE, openWorldHint: true, // Accepts any user-supplied URL on the public web. destructiveHint: false, // Does not modify, delete, or write to external websites. }, @@ -2482,6 +2535,9 @@ Alexandria mode, on an authenticated session with Alexandria access: \`alexandri } & Record; if (alexandria) { assertExchangeCredential(session); + if ((Array.isArray(alexandria) ? alexandria : [alexandria]).some(isTermsWrite)) { + throw termsWriteError(); + } log.info('Executing Alexandria capabilities', { count: Array.isArray(alexandria) ? alexandria.length : 1, }); @@ -2492,7 +2548,7 @@ Alexandria mode, on an authenticated session with Alexandria access: \`alexandri const transformed = transformScrapeParams( options as Record ); - const cleaned = removeEmptyTopLevel(transformed); + const cleaned = withReadOnlyProfile(removeEmptyTopLevel(transformed)); if (cleaned.lockdown) { log.info('Scraping URL (lockdown)'); } else { @@ -2592,6 +2648,7 @@ For a programming question, add \`categories: ["developer"]\`; its hits return i ...searchToolBaseFields, scrapeOptions: scrapeParamsSchema .omit({ url: true }) + .extend(SAFE_MODE ? { profile: readOnlyProfileSchema } : {}) .partial() .optional() .describe('Attach page content for web results in the same call. These fetches ignore maxAge, so use firecrawl_scrape when you need a live fetch. scrapeOptions fetches web pages, never Alexandria provider tools.'), @@ -2614,8 +2671,8 @@ For a programming question, add \`categories: ["developer"]\`; its hits return i searchOpts.toolDetail ??= 'compact'; if (searchOpts.scrapeOptions) { - searchOpts.scrapeOptions = transformScrapeParams( - searchOpts.scrapeOptions as Record + searchOpts.scrapeOptions = withReadOnlyProfile( + transformScrapeParams(searchOpts.scrapeOptions as Record) ); } @@ -2753,13 +2810,14 @@ const findToolsTool: RegisteredTool = { }, }; server.addTool(findToolsTool); + // Search-surface copies of the two Alexandria tools. // Same parameters and executor as the full-surface tools; only the // descriptions differ, so they name nothing that surface does not register // (no crawl, map, interact, monitor, parse or feedback references). Registered // on the search surface in place of the module-level tools above. const SEARCH_SURFACE_SCRAPE_DESCRIPTION = ` -Scrape one URL and return its content, or execute catalogued Alexandria capabilities. URL mode returns markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema, plus page metadata. Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch. A successful response does not by itself confirm the page is still current. Browser actions can change the live page when interactive actions are enabled, and a named browser profile can load saved session data and overwrite its stored state. +Scrape one URL and return its content, or execute catalogued Alexandria capabilities. URL mode returns markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema, plus page metadata. Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch. A successful response does not by itself confirm the page is still current. ${SAFE_MODE ? 'A named browser profile loads saved session data without saving changes to it.' : 'Browser actions can change the live page, and a named browser profile can load saved session data and overwrite its stored state.'} \`firecrawl_search\` with \`sources\` unset and \`firecrawl_find_tools\` can discover providers for the same fields across several pages; a matching Alexandria provider returns typed records in one call. @@ -3519,7 +3577,7 @@ server.addTool({ name: 'firecrawl_agent', annotations: { title: 'Firecrawl agent', - readOnlyHint: false, // Starts an autonomous research agent job on the Firecrawl API. + readOnlyHint: false, // The research agent can click, fill forms, and navigate interactive pages. openWorldHint: true, // The agent browses and searches the open web to fulfill the prompt. destructiveHint: false, // Gathers information only; does not delete external data or user resources. }, diff --git a/src/instructions.ts b/src/instructions.ts new file mode 100644 index 00000000..349b7690 --- /dev/null +++ b/src/instructions.ts @@ -0,0 +1,33 @@ +// Server instructions, one string per tool surface. Clients read them before a +// tool is loaded: Claude Code truncates server instructions at 2,048 +// characters, OpenAI asks for the key details in the first 512, and Codex code +// mode prepends them to every tool entry in ALL_TOOLS. Each string routes +// between the tools its surface registers and stays under 1,024 characters; +// tests/mcp-instructions.test.mjs checks both against what each surface +// actually serves (budgets in tests/helpers/instructions.mjs). + +/** Keyed and OAuth sessions on the full surface (hosted /v2/mcp, /v2/mcp-oauth, stdio with a key). */ +export const FULL_INSTRUCTIONS = + 'Firecrawl gives agents live web data: search the web, read pages, and collect data across sites. ' + + 'Pick the tool by what you have: no URL, firecrawl_search; one known URL, firecrawl_scrape (formats: ["json"] for fields); ' + + "a site's URLs, firecrawl_map; many pages of one site, firecrawl_crawl; pages behind clicks, forms, or login, firecrawl_interact; " + + 'URLs unknown or the answer spans many sites, firecrawl_agent; a local file, firecrawl_parse. ' + + 'Programming questions: firecrawl_developer_search. Research papers: firecrawl_research_search_papers, then the other firecrawl_research_* tools. ' + + 'Recurring page checks: firecrawl_monitor_*. ' + + 'firecrawl_agent returns a job ID; read the result with firecrawl_agent_status. ' + + "firecrawl_search also returns matching Alexandria data providers in data.tools; firecrawl_find_tools reads a provider's contract and firecrawl_scrape with alexandria runs it."; + +/** The search surface (/v2/mcp-search). */ +export const SEARCH_INSTRUCTIONS = + 'Firecrawl Search: web, developer, and research search, plus Alexandria data providers. ' + + 'No URL, firecrawl_search; one known URL, firecrawl_scrape with url. ' + + 'Programming questions (code, libraries, APIs, errors): firecrawl_developer_search. ' + + 'Research papers: firecrawl_research_search_papers, then firecrawl_research_inspect_paper, firecrawl_research_related_papers, or firecrawl_research_read_paper. ' + + "firecrawl_search also returns matching Alexandria providers in data.tools; firecrawl_find_tools reads a provider's contract and firecrawl_scrape with alexandria runs it. " + + 'Web, developer, and research searches and URL scrapes are billed per request, Alexandria capabilities at their listed price; provider discovery is free.'; + +/** Keyless sessions (hosted keyless tier, stdio without a key). */ +export const KEYLESS_INSTRUCTIONS = + 'Firecrawl keyless access is usage-limited. No URL, firecrawl_search (categories: ["developer"] for programming questions); ' + + 'one known URL, firecrawl_scrape; a local file, firecrawl_parse. ' + + 'An API key adds site mapping, crawling, interaction, and the research agent, with higher limits.'; diff --git a/tests/helpers/alexandria-metadata.mjs b/tests/helpers/alexandria-metadata.mjs index ba17fa13..3c6170c1 100644 --- a/tests/helpers/alexandria-metadata.mjs +++ b/tests/helpers/alexandria-metadata.mjs @@ -1,13 +1,15 @@ import assert from 'node:assert/strict'; -export function assertAlexandriaMetadata(tools, instructions) { +export function assertAlexandriaMetadata(tools, instructions, { hosted }) { assert.ok(instructions, 'Alexandria server instructions are required'); const scrape = tools.find((tool) => tool.name === 'firecrawl_scrape'); assert.ok(scrape, 'firecrawl_scrape must be registered'); const { requestId, alexandria } = scrape.inputSchema?.properties ?? {}; assert.ok(requestId?.description, 'firecrawl_scrape.requestId needs a description'); assert.ok(alexandria?.description, 'firecrawl_scrape.alexandria needs a description'); - assert.equal(scrape.annotations.readOnlyHint, false); + // Hosted (safe mode) scrape has no browser actions, loads profiles + // read-only, and refuses terms acceptance, which happens in the dashboard. + assert.equal(scrape.annotations.readOnlyHint, hosted); assert.match(requestId.description, /idempotency key.*payload/i); assert.match(requestId.description, /generated when omitted and returned with the result/i); assert.match(alexandria.description, /mutually exclusive with url/); diff --git a/tests/helpers/exchange-mcp.mjs b/tests/helpers/exchange-mcp.mjs index 578062ee..7ce98fd8 100644 --- a/tests/helpers/exchange-mcp.mjs +++ b/tests/helpers/exchange-mcp.mjs @@ -174,6 +174,7 @@ async function startStdioWithApi(t, options = {}) { const api = await startFakeExchangeApi(options); t.after(() => api.close()); const session = await startStdio(t, { + CLOUD_SERVICE: 'false', FIRECRAWL_API_KEY: 'fc-exchange-test', FIRECRAWL_API_URL: api.url, }); diff --git a/tests/helpers/instructions.mjs b/tests/helpers/instructions.mjs new file mode 100644 index 00000000..702d72a3 --- /dev/null +++ b/tests/helpers/instructions.mjs @@ -0,0 +1,30 @@ +import assert from 'node:assert/strict'; + +// The instructions budget every surface must meet (see src/instructions.ts). +export const INSTRUCTIONS_MAX_CHARS = 1024; +// OpenAI asks for the key details of server instructions in the first 512 characters. +export const ROUTING_HEAD_CHARS = 512; + +/** + * Server instructions fit the cap, route search and scrape inside the first + * 512 characters, and name only tools this session lists. A trailing `*` + * names a family and needs at least one listed member. + */ +export function assertInstructionsMatchTools(instructions, tools, label) { + assert.ok(instructions, `${label}: instructions are required`); + assert.ok( + instructions.length <= INSTRUCTIONS_MAX_CHARS, + `${label}: instructions are ${instructions.length} characters` + ); + const head = instructions.slice(0, ROUTING_HEAD_CHARS); + for (const name of ['firecrawl_search', 'firecrawl_scrape']) { + assert.ok(head.includes(name), `${label}: ${name} is routed in the first ${ROUTING_HEAD_CHARS} characters`); + } + const listed = tools.map((tool) => tool.name); + for (const token of new Set(instructions.match(/firecrawl_[a-z_]+\*?/g) ?? [])) { + const present = token.endsWith('*') + ? listed.some((name) => name.startsWith(token.slice(0, -1))) + : listed.includes(token); + assert.ok(present, `${label}: instructions name ${token}, which this session does not list`); + } +} diff --git a/tests/mcp-alexandria-terms.test.mjs b/tests/mcp-alexandria-terms.test.mjs index 4ec1ec07..82e87b19 100644 --- a/tests/mcp-alexandria-terms.test.mjs +++ b/tests/mcp-alexandria-terms.test.mjs @@ -3,16 +3,19 @@ import test from 'node:test'; import { TERMS_REQUIRED_BODY } from './helpers/exchange-api.mjs'; import { startStdioWithApi, callExpectingError, toolText } from './helpers/exchange-mcp.mjs'; -test('terms are disclosed after a blocked provider and use scrape instead of top-level tools', async (t) => { +test('terms are read through scrape and accepted by an organization admin in the dashboard', async (t) => { const { api, client } = await startStdioWithApi(t); const listing = await client.request('tools/list', {}); - assert.ok(!listing.tools.some(tool => /^firecrawl_terms_/.test(tool.name))); + assert.ok(!listing.tools.some(tool => /terms/.test(tool.name)), 'no tool accepts terms'); + const blocked = await callExpectingError(client, { name: 'firecrawl_scrape', arguments: { alexandria: [{ provider: 'benzinga', capability: 'news/search' }] } }); assert.equal(api.requests.length, 1, 'blocked requests never auto-accept'); - assert.match(blocked.content[0].text, /explicit authorization to accept that exact version and digest/); + const { code, status, requiresAction, requestId, next_actions } = blocked.structuredContent; + assert.match(blocked.content[0].text, /Terms are accepted in the Firecrawl dashboard, not through this connection: ask an organization admin to accept them at /); + assert.ok(blocked.content[0].text.includes(requiresAction.url), 'names the admin accept page'); assert.match(blocked.content[0].text, /Never infer acceptance from a data request/); assert.match(blocked.content[0].text, /Reuse this ID only with the identical payload/); - const { code, status, requiresAction, requestId, next_actions } = blocked.structuredContent; + assert.doesNotMatch(blocked.content[0].text, /terms\/accept|confirmed: ?true/); assert.equal(code, TERMS_REQUIRED_BODY.code); assert.equal(status, 403); assert.deepEqual(requiresAction, TERMS_REQUIRED_BODY.requiresAction); @@ -30,13 +33,18 @@ test('terms are disclosed after a blocked provider and use scrape instead of top assert.equal(next.arguments.alexandria[0].capability, 'terms/show'); const read = toolText(await client.request('tools/call', next)).data.alexandria[0].data; assert.equal(read.terms.document, 'Review this agreement.'); + assert.equal(api.requests.length, 2); + const options = { provider: 'benzinga', version: read.terms.version, digest: read.terms.digest, confirmed: true }; - const accepted = toolText(await client.request('tools/call', { name: 'firecrawl_scrape', arguments: { alexandria: [{ provider: 'firecrawl', capability: 'terms/accept', options }] } })); - assert.ok(accepted.data.alexandria[0].data.acceptedAt); - assert.equal(api.requests.length, 3); - assert.ok(api.requests.every(request => request.url === '/v2/scrape')); - assert.deepEqual(api.requests[2].body.alexandria[0].options, options); - assert.equal(api.requests[2].headers.authorization, 'Bearer fc-exchange-test'); + for (const call of [ + { provider: 'firecrawl', capability: 'terms/accept', options }, + { provider: ' Firecrawl ', capability: 'Terms/Accept', options }, + { provider: 'firecrawl', capability: 'terms/revoke', options: { provider: 'benzinga' } }, + ]) { + const refused = await callExpectingError(client, { name: 'firecrawl_scrape', arguments: { alexandria: [call] } }); + assert.match(refused.content[0].text, /Provider terms are accepted in the Firecrawl dashboard, not through this connection/, call.capability); + } + assert.equal(api.requests.length, 2, 'scrape forwards no terms command other than terms/show'); }); test('firecrawl_scrape relays a reserved 409 billing error with its code and chargeId', async (t) => { @@ -76,7 +84,8 @@ test('disabled provider offers read-only terms recovery without treating other r assert.equal(blocked.structuredContent.status, 403); assert.equal(Boolean(blocked.structuredContent.nextTool), expected, message); if (expected) { - assert.match(blocked.content[0].text, /explicit authorization/); + assert.match(blocked.content[0].text, /Terms are accepted in the Firecrawl dashboard, not through this connection/); + assert.doesNotMatch(blocked.content[0].text, /terms\/accept/); const shown = toolText(await client.request('tools/call', blocked.structuredContent.nextTool)); assert.equal(shown.data.alexandria[0].data.status.accepted, false); assert.equal(api.requests.length, 2); diff --git a/tests/mcp-description-budget.test.mjs b/tests/mcp-description-budget.test.mjs index fe8da66e..c1687e07 100644 --- a/tests/mcp-description-budget.test.mjs +++ b/tests/mcp-description-budget.test.mjs @@ -12,7 +12,7 @@ import { assertAlexandriaMetadata } from './helpers/alexandria-metadata.mjs'; test('every tool description fits the 2,048-character cap and keeps the routing copy inside it', async (t) => { const { client, init } = await startStdioWithApi(t); const { tools } = await client.request('tools/list', {}); - assertAlexandriaMetadata(tools, init.instructions); + assertAlexandriaMetadata(tools, init.instructions, { hosted: false }); for (const tool of tools) { assert.ok((tool.description ?? '').length <= CAP, `${tool.name} description is ${(tool.description ?? '').length} chars`); } diff --git a/tests/mcp-instructions.test.mjs b/tests/mcp-instructions.test.mjs new file mode 100644 index 00000000..089f3da0 --- /dev/null +++ b/tests/mcp-instructions.test.mjs @@ -0,0 +1,194 @@ +import assert from 'node:assert/strict'; +import test from 'node:test'; +import { startFakeExchangeApi } from './helpers/exchange-api.mjs'; +import { + getFreePort, + parseSseJson, + spawnServer, + startStdio, + stopChild, + waitForHealth, +} from './helpers/exchange-mcp.mjs'; +import { assertInstructionsMatchTools } from './helpers/instructions.mjs'; + +const FULL_PREFIX = /^Firecrawl gives agents live web data/; +const KEYLESS_PREFIX = /^Firecrawl keyless access is usage-limited/; +const SEARCH_PREFIX = /^Firecrawl Search: web, developer, and research search/; + +function rpc(port, endpoint, { id, method, params = {}, headers = {} }) { + return fetch(`http://127.0.0.1:${port}${endpoint}`, { + body: JSON.stringify({ id, jsonrpc: '2.0', method, params }), + headers: { + accept: 'application/json, text/event-stream', + 'content-type': 'application/json', + ...headers, + }, + method: 'POST', + }); +} + +async function rpcResult(port, endpoint, request) { + const response = await rpc(port, endpoint, request); + assert.equal(response.status, 200, `${request.method} returned ${response.status}`); + return parseSseJson(await response.text()).result; +} + +async function httpSession(port, endpoint, headers) { + const init = await rpcResult(port, endpoint, { + id: 1, + method: 'initialize', + params: { + capabilities: {}, + clientInfo: { name: 'firecrawl-instructions-test', version: '0.0.0' }, + protocolVersion: '2025-06-18', + }, + headers, + }); + const { tools } = await rpcResult(port, endpoint, { id: 2, method: 'tools/list', headers }); + return { instructions: init.instructions, tools }; +} + +async function startHosted(t) { + const api = await startFakeExchangeApi(); + t.after(() => api.close()); + const port = await getFreePort(); + const searchPort = await getFreePort(); + const child = spawnServer({ + CLOUD_SERVICE: 'true', + HTTP_STREAMABLE_SERVER: 'true', + FASTMCP_ENDPOINT: '/v2/mcp', + FIRECRAWL_API_KEY: '', + FIRECRAWL_OAUTH_TOKEN: '', + FIRECRAWL_API_URL: api.url, + KEYLESS_PROXY_SECRET: 'keyless-secret', + PORT: String(port), + FIRECRAWL_MCP_SEARCH_PORT: String(searchPort), + }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + await waitForHealth(searchPort, child); + return { api, port, searchPort }; +} + +test('stdio sessions get the instructions for the tools they can use', async (t) => { + const api = await startFakeExchangeApi(); + t.after(() => api.close()); + for (const [label, env, prefix] of [ + ['stdio with an API key', { FIRECRAWL_API_KEY: 'fc-test', FIRECRAWL_API_URL: '' }, FULL_PREFIX], + ['stdio self-hosted', { FIRECRAWL_API_KEY: '', FIRECRAWL_OAUTH_TOKEN: '', FIRECRAWL_API_URL: api.url }, FULL_PREFIX], + ['stdio keyless', { FIRECRAWL_API_KEY: '', FIRECRAWL_OAUTH_TOKEN: '', FIRECRAWL_API_URL: '' }, KEYLESS_PREFIX], + ]) { + const { client, init } = await startStdio(t, { CLOUD_SERVICE: 'false', ...env }); + const { tools } = await client.request('tools/list', {}); + assert.match(init.instructions, prefix, label); + assertInstructionsMatchTools(init.instructions, tools, label); + } +}); + +test('local HTTP sessions, which always carry a key or API URL, get the full instructions', async (t) => { + const port = await getFreePort(); + const child = spawnServer({ + HTTP_STREAMABLE_SERVER: 'true', + HOST: '127.0.0.1', + CLOUD_SERVICE: 'false', + FIRECRAWL_API_KEY: '', + FIRECRAWL_OAUTH_TOKEN: '', + FIRECRAWL_API_URL: '', + PORT: String(port), + }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + for (const [label, headers] of [ + ['local HTTP with x-api-key', { 'x-api-key': 'fc-local-test' }], + ['local HTTP with a bearer key', { authorization: 'Bearer fc-local-test' }], + ]) { + const { instructions, tools } = await httpSession(port, '/mcp', headers); + assert.match(instructions, FULL_PREFIX, label); + assertInstructionsMatchTools(instructions, tools, label); + } +}); + +test('hosted sessions get instructions that match their own tool list', async (t) => { + const { port, searchPort } = await startHosted(t); + for (const [label, headers, prefix] of [ + ['hosted keyless', { 'x-forwarded-for': '8.8.8.8' }, KEYLESS_PREFIX], + ['hosted with an API key', { 'x-api-key': 'fc-hosted-test' }, FULL_PREFIX], + ['hosted with a malformed credential', { authorization: 'Bearer not-a-firecrawl-credential' }, KEYLESS_PREFIX], + ]) { + const { instructions, tools } = await httpSession(port, '/v2/mcp', headers); + assert.match(instructions, prefix, label); + assertInstructionsMatchTools(instructions, tools, label); + } + + const search = await httpSession(searchPort, '/v2/mcp-search', { 'x-api-key': 'fc-hosted-test' }); + assert.match(search.instructions, SEARCH_PREFIX); + assertInstructionsMatchTools(search.instructions, search.tools, 'search surface'); +}); + +test('hosted scrape is read-only; local scrape and the research agent are not', async (t) => { + const { api, port, searchPort } = await startHosted(t); + const headers = { 'x-api-key': 'fc-hosted-test' }; + for (const [label, endpoint, surfacePort] of [ + ['hosted full surface', '/v2/mcp', port], + ['search surface', '/v2/mcp-search', searchPort], + ]) { + const { tools } = await httpSession(surfacePort, endpoint, headers); + const scrape = tools.find((tool) => tool.name === 'firecrawl_scrape'); + assert.equal(scrape.annotations.readOnlyHint, true, label); + assert.equal(scrape.annotations.destructiveHint, false, label); + assert.equal(scrape.inputSchema.properties.actions, undefined, `${label}: no browser actions`); + assert.deepEqual(Object.keys(scrape.inputSchema.properties.profile.properties), ['name'], `${label}: profiles load read-only`); + assert.doesNotMatch(scrape.description, /overwrite its stored state/, label); + } + + const { tools } = await httpSession(port, '/v2/mcp', headers); + const byName = new Map(tools.map((tool) => [tool.name, tool])); + assert.equal(byName.get('firecrawl_agent').annotations.readOnlyHint, false); + assert.deepEqual( + Object.keys(byName.get('firecrawl_search').inputSchema.properties.scrapeOptions.properties.profile.properties), + ['name'] + ); + assert.equal(tools.some((tool) => /terms/.test(tool.name)), false, 'no tool accepts terms'); + // Saving a profile stays available through interact. + assert.ok(byName.get('firecrawl_interact').inputSchema.properties.scrapeOptions.properties.profile.properties.saveChanges); + + const scraped = await rpcResult(port, '/v2/mcp', { + id: 3, + method: 'tools/call', + params: { + name: 'firecrawl_scrape', + arguments: { url: 'https://example.com/account', profile: { name: 'saved-login', saveChanges: true } }, + }, + headers, + }); + assert.notEqual(scraped.isError, true, JSON.stringify(scraped)); + const sent = api.requests.find((request) => request.url === '/v2/scrape' && request.body?.url); + assert.deepEqual(sent.body.profile, { name: 'saved-login', saveChanges: false }); + + const { client } = await startStdio(t, { CLOUD_SERVICE: 'false', FIRECRAWL_API_KEY: 'fc-test', FIRECRAWL_API_URL: api.url }); + const local = (await client.request('tools/list', {})).tools.find((tool) => tool.name === 'firecrawl_scrape'); + assert.equal(local.annotations.readOnlyHint, false); + assert.ok(local.inputSchema.properties.actions, 'local scrape keeps browser actions'); + assert.ok(local.inputSchema.properties.profile.properties.saveChanges, 'local scrape keeps writable profiles'); +}); + +test('hosted scrape refuses terms acceptance on both surfaces and points to the dashboard', async (t) => { + const { api, port, searchPort } = await startHosted(t); + const before = api.requests.length; + for (const [surfacePort, endpoint] of [[port, '/v2/mcp'], [searchPort, '/v2/mcp-search']]) { + const result = await rpcResult(surfacePort, endpoint, { + id: 4, + method: 'tools/call', + params: { + name: 'firecrawl_scrape', + arguments: { + alexandria: [{ provider: 'firecrawl', capability: 'terms/accept', options: { provider: 'benzinga', version: 'v1', digest: 'a'.repeat(64), confirmed: true } }], + }, + }, + headers: { 'x-api-key': 'fc-hosted-test' }, + }); + assert.equal(result.isError, true, endpoint); + assert.match(result.content[0].text, /accepted in the Firecrawl dashboard, not through this connection.*https:\/\/www\.firecrawl\.dev\/app\/settings\?tab=data-sources/, endpoint); + } + assert.equal(api.requests.length, before, 'terms/accept never reaches the API through scrape'); +}); diff --git a/tests/mcp-search-profile.test.mjs b/tests/mcp-search-profile.test.mjs index 2efe5382..8efa1335 100644 --- a/tests/mcp-search-profile.test.mjs +++ b/tests/mcp-search-profile.test.mjs @@ -8,6 +8,7 @@ import { setTimeout as delay } from 'node:timers/promises'; import { assertAgentMetadataPolicy } from '../scripts/agent-metadata-policy.mjs'; import { CLAUDE_CODE_TEXT_CAP } from './helpers/description-budget.mjs'; import { assertAlexandriaMetadata } from './helpers/alexandria-metadata.mjs'; +import { assertInstructionsMatchTools } from './helpers/instructions.mjs'; const { version: serverVersion } = JSON.parse( readFileSync(new URL('../package.json', import.meta.url), 'utf8') @@ -1066,7 +1067,8 @@ test('primary search profile agent language satisfies metadata policy gates', as const initialize = await initializeProfile(port, SEARCH_ENDPOINT, headers); const tools = await listToolDefinitions(port, SEARCH_ENDPOINT, headers); - assertAlexandriaMetadata(tools, initialize.instructions); + assertAlexandriaMetadata(tools, initialize.instructions, { hosted: true }); + assertInstructionsMatchTools(initialize.instructions, tools, 'hosted profile'); assertAgentMetadataPolicy( [initialize.instructions, ...tools.map((tool) => tool.description ?? '')], assert @@ -1074,14 +1076,17 @@ test('primary search profile agent language satisfies metadata policy gates', as }); test('keyless full-surface instructions satisfy the same metadata policy gates', async (t) => { - // FASTMCP_ENDPOINT '/v2/mcp' (not '/v2/mcp-oauth') makes this the keyless - // profile, whose instructions must stay descriptive rather than becoming an - // imperative routing playbook. + // FASTMCP_ENDPOINT '/v2/mcp' (not '/v2/mcp-oauth') allows keyless + // sessions; a request without a credential gets the keyless instructions, + // whose wording must stay descriptive rather than becoming an imperative + // routing playbook. const { fullPort } = await startHostedServer(t); - const headers = { 'x-api-key': 'fc-keyless-metadata' }; + const headers = { 'x-forwarded-for': '8.8.8.8' }; const initialize = await initializeProfile(fullPort, '/v2/mcp', headers); const tools = await listToolDefinitions(fullPort, '/v2/mcp', headers); + assert.match(initialize.instructions, /^Firecrawl keyless access is usage-limited/); + assertInstructionsMatchTools(initialize.instructions, tools, 'hosted keyless'); assertAgentMetadataPolicy( [initialize.instructions, ...tools.map((tool) => tool.description ?? '')], assert @@ -1114,7 +1119,8 @@ test('account (mcp-oauth) full-surface instructions satisfy the same metadata po assert.equal(tool?._meta?.['anthropic/alwaysLoad'], true, name); } - assertAlexandriaMetadata(tools, initialize.instructions); + assertAlexandriaMetadata(tools, initialize.instructions, { hosted: true }); + assertInstructionsMatchTools(initialize.instructions, tools, 'hosted profile'); assertAgentMetadataPolicy( [initialize.instructions, ...tools.map((tool) => tool.description ?? '')], assert @@ -1309,7 +1315,7 @@ test('search surface registers the two Alexandria tools with surface-scoped copy const headers = { 'x-api-key': 'fc-test' }; const initialize = await initializeProfile(searchPort, SEARCH_ENDPOINT, headers); const tools = await listToolDefinitions(searchPort, SEARCH_ENDPOINT, headers); - assertAlexandriaMetadata(tools, initialize.instructions); + assertAlexandriaMetadata(tools, initialize.instructions, { hosted: true }); // Claude Code truncates tool descriptions at CLAUDE_CODE_TEXT_CAP characters. for (const tool of tools) { assert.ok((tool.description ?? '').length <= CLAUDE_CODE_TEXT_CAP, `${tool.name} description is ${(tool.description ?? '').length} chars`); diff --git a/tests/mcp-smoke.test.mjs b/tests/mcp-smoke.test.mjs index 21abe6ff..e45d76c2 100644 --- a/tests/mcp-smoke.test.mjs +++ b/tests/mcp-smoke.test.mjs @@ -7,6 +7,7 @@ import test from 'node:test'; import { setTimeout as delay } from 'node:timers/promises'; import { assertAgentMetadataPolicy } from '../scripts/agent-metadata-policy.mjs'; import { CLAUDE_CODE_TEXT_CAP } from './helpers/description-budget.mjs'; +import { ROUTING_HEAD_CHARS } from './helpers/instructions.mjs'; const { version: serverVersion } = JSON.parse( readFileSync(new URL('../package.json', import.meta.url), 'utf8') @@ -852,9 +853,10 @@ test('HTTP cloud keyless transport preserves app challenge without advertising O assert.match(initialize.headers.get('content-type') ?? '', /text\/event-stream/); const initializeMessage = parseSseJson(await initialize.text()); assert.equal(initializeMessage.result.serverInfo.name, 'firecrawl-fastmcp'); + // A keyed session lists the full tool set, so it gets the full instructions. assert.match( initializeMessage.result.instructions, - /An Authorization bearer API key can provide higher usage limits and expose additional tools/i + /^Firecrawl gives agents live web data.*firecrawl_map/ ); assert.doesNotMatch(initializeMessage.result.instructions, /\bOAuth\b/i); @@ -1158,26 +1160,17 @@ test('stdio transport initializes and lists Firecrawl tools', async (t) => { .description, 'Break historical usage down by API key. When view is omitted, true selects the historical view; it cannot be combined with view "current".' ); - // A stdio session with an API key gets the Alexandria-aware instructions, - // not the keyless wording. - assert.match(init.instructions, /firecrawl_scrape retrieves one supplied page/i); - // Keep Alexandria routing within the truncated server instructions. The search - // description also carries developer and research routing (asserted below). - const instructionsHead = init.instructions.slice(0, CLAUDE_CODE_TEXT_CAP); - assert.match(instructionsHead, /Alexandria is Firecrawl's catalogue of data providers/); - assert.match(instructionsHead, /For the same fields across multiple pages/); - assert.match(instructionsHead, /sources: \["web"\] omits semantic provider discovery/); - assert.match( - init.instructions, - /Alexandria is Firecrawl's catalogue of data providers and workflows.*firecrawl_scrape with alexandria.*executes up to ten capabilities/is - ); - assert.match( - init.instructions, - /sources: \["web"\] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false/ - ); + // A stdio session with an API key gets the full instructions, not the + // keyless wording, and the routing map sits inside OpenAI's first 512 + // characters. + assert.match(init.instructions, /^Firecrawl gives agents live web data/); + const routingHead = init.instructions.slice(0, ROUTING_HEAD_CHARS); + for (const name of ['firecrawl_search', 'firecrawl_scrape', 'firecrawl_map', 'firecrawl_crawl', 'firecrawl_interact', 'firecrawl_agent', 'firecrawl_parse']) { + assert.ok(routingHead.includes(name), `${name} is routed in the first ${ROUTING_HEAD_CHARS} characters`); + } assert.match( init.instructions, - /For the same fields across multiple pages, firecrawl_find_tools offers free provider discovery/ + /firecrawl_search also returns matching Alexandria data providers in data\.tools; firecrawl_find_tools reads a provider's contract and firecrawl_scrape with alexandria runs it/ ); assert.match( byName.get('firecrawl_scrape').description, @@ -1289,12 +1282,8 @@ test('stdio transport initializes and lists Firecrawl tools', async (t) => { /data\.developer/i ); assert.match( - init.instructions, - /firecrawl_search with categories: \["research"\] is a website filter over ordinary web results and reaches different sources/i - ); - assert.match( - init.instructions, - /firecrawl_research_\* tools search a paper index of abstracts and full text/i + byName.get('firecrawl_search').description, + /`categories: \["research"\]` restricts web results to research-affiliated websites; the `firecrawl_research_\*` tools are a separate surface over paper abstracts and full text/ ); assert.match( byName.get('firecrawl_research_related_papers').description, @@ -1422,42 +1411,16 @@ test('local keyless stdio keeps profile guidance keyless-scoped and omits feedba clientInfo: { name: 'firecrawl-local-keyless', version: '0.0.0' }, protocolVersion: '2025-06-18', }); - const apiKeyBoundary = 'An Authorization bearer API key'; - const apiKeyBoundaryIndex = init.instructions.indexOf(apiKeyBoundary); - assert.notEqual(apiKeyBoundaryIndex, -1); - const keylessGuidance = init.instructions.slice(0, apiKeyBoundaryIndex); - const apiKeyGuidance = init.instructions.slice(apiKeyBoundaryIndex); - assert.match( - keylessGuidance, - /Hosted keyless sessions expose firecrawl_search, firecrawl_scrape, and firecrawl_parse with usage limits/i - ); - assert.match( - keylessGuidance, - /firecrawl_search with categories: \["developer"\].*code documentation/i - ); - assert.match( - keylessGuidance, - /firecrawl_search with categories: \["research"\].*research-affiliated websites/i - ); - assert.match(keylessGuidance, /firecrawl_scrape retrieves one supplied page/i); - assert.match(keylessGuidance, /firecrawl_parse processes supported local files/i); - assert.doesNotMatch( - keylessGuidance, - /firecrawl_(?:map|agent|agent_status|research_)/i + // Keyless instructions name only the keyless tools and describe what an + // API key adds without naming tools this session cannot call. + assert.match(init.instructions, /^Firecrawl keyless access is usage-limited/); + assert.deepEqual( + [...new Set(init.instructions.match(/firecrawl_[a-z_*]+/g))].sort(), + ['firecrawl_parse', 'firecrawl_scrape', 'firecrawl_search'] ); + assert.match(init.instructions, /firecrawl_search \(categories: \["developer"\] for programming questions\)/); + assert.match(init.instructions, /An API key adds site mapping, crawling, interaction, and the research agent, with higher limits/); assert.doesNotMatch(init.instructions, /\bOAuth\b/i); - assert.match( - apiKeyGuidance, - /higher usage limits.*additional tools.*subject to plan, deployment, and team policy/is - ); - for (const name of [ - 'firecrawl_map', - 'firecrawl_agent', - 'firecrawl_agent_status', - 'firecrawl_research_*', - ]) { - assert.ok(apiKeyGuidance.includes(name), `${name} must be API-key qualified`); - } client.notify('notifications/initialized'); const tools = await client.request('tools/list'); const toolNames = tools.tools.map((tool) => tool.name); From 2c61c534594f91e0929e692c5af38faa2ff01b53 Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Wed, 30 Sep 2026 22:11:33 -0500 Subject: [PATCH 2/3] fix(mcp): route research agent terms acceptance to dashboard --- CHANGELOG.md | 2 +- README.md | 14 +++++++------- src/index.ts | 10 +++++----- src/tool-output.ts | 4 ++-- tests/mcp-smoke.test.mjs | 22 +++++++++++++++------- 5 files changed, 30 insertions(+), 22 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 00815714..02ab5f6f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,7 +6,7 @@ - `firecrawl_agent` now exposes the optional `effort` (`low`, `medium`, `high`), `maxCredits`, and `strictConstrainToURLs` parameters that `POST /v2/agent` already accepts, and forwards them in the request body. - `firecrawl_agent` can continue a thread: it accepts `threadId` and `mode` (`"extract"` or `"chat"`) and forwards them to `POST /v2/agent` through the SDK. On a follow-up, omitted `mode`, `urls` and `schema` carry over from the previous turn. `firecrawl_agent_status` now keeps `message` and `suggestions` in its structured content, next to `threadId` and `threadTurn`. -- `firecrawl_agent` accepts an `exchange` object that mirrors the API's (`enabled`, `toolkits` (at most 5), `maxCalls`, `requireApproval`, `approve: { approvalId, callIds?, always? }`, `decline: { approvalId }`, `onTermsRequired`) and forwards it to `POST /v2/agent` through the SDK. `exchange.onTermsRequired` (`"skip"` or `"ask"`) controls Alexandria providers whose data terms the team has not accepted; they are never called. After an ask-mode terms offer and the user's explicit consent to `terms/accept`, a caller answers the offer on the same thread with `exchange.approve: { approvalId }` (or `exchange.decline`) instead of starting over. `approve` and `decline` require `threadId` and cannot be sent together, and `exchange.requireApproval` requires `mode: "chat"` on the same call. There is no auto-accept. `firecrawl_agent_status` now keeps `exchange` (including `skippedProviders` and `requiresAction`, whose provider `digest` is `string | null` and always present) and `pendingApproval` in its structured content. +- `firecrawl_agent` accepts an `exchange` object that mirrors the API's (`enabled`, `toolkits` (at most 5), `maxCalls`, `requireApproval`, `approve: { approvalId, callIds?, always? }`, `decline: { approvalId }`, `onTermsRequired`) and forwards it to `POST /v2/agent` through the SDK. `exchange.onTermsRequired` (`"skip"` or `"ask"`) controls Alexandria providers whose data terms the team has not accepted; they are never called. After an ask-mode terms offer and an organization admin's confirmed dashboard acceptance, a caller resumes the same thread with `exchange.approve: { approvalId }` (or declines with `exchange.decline`) instead of starting over. Approval does not accept terms. `approve` and `decline` require `threadId` and cannot be sent together, and `exchange.requireApproval` requires `mode: "chat"` on the same call. There is no auto-accept. `firecrawl_agent_status` now keeps `exchange` (including `skippedProviders` and `requiresAction`, whose provider `digest` is `string | null` and always present) and `pendingApproval` in its structured content. ### Changed diff --git a/README.md b/README.md index 05435717..75bdfba2 100644 --- a/README.md +++ b/README.md @@ -748,17 +748,17 @@ The agent performs web searches, follows links, reads pages, and gathers data au - `enabled`, `toolkits` (up to 5 provider slugs), `maxCalls` (1 to 30), `requireApproval` (paid calls end the turn with a `pendingApproval`; needs `mode: "chat"` on the same call, even on a follow-up) - `onTermsRequired`: what to do when an Alexandria provider the agent would use needs data terms your team has not accepted. Gated providers are never called in any mode. Omitted on a follow-up keeps the previous turn's value. - `"skip"` (default): answer with accepted providers only. `exchange.skippedProviders` on the status result lists the gated providers that would have helped. - - `"ask"`: the same, plus a terms `pendingApproval` and `exchange.requiresAction` with the exact `terms/show` and `terms/accept` calls for each provider. Each provider's `digest` is always present and is `string | null`; when it is `null`, `terms/show` returns the current digest to send. + - `"ask"`: the same, plus a terms `pendingApproval` and `exchange.requiresAction` with the approval ID and provider requirements. Read terms with `terms/show`; an organization admin accepts them in the Firecrawl dashboard. - `approve`: `{ approvalId, callIds?, always? }` answers yes to the `pendingApproval` the previous turn ended on. `callIds` and `always` apply to paid-call approvals only. - `decline`: `{ approvalId }` answers no. A declined terms offer keeps those providers out of the rest of the thread. - `approve` and `decline` need `threadId`, and only one of them can be sent. -**Provider terms (ask mode):** there is no auto-accept mode. When a turn ends on a terms offer, the status result carries `pendingApproval` (`kind: "terms"`) and `exchange.requiresAction` with the `approvalId` and the exact `terms/show` and `terms/accept` calls. To use the provider: +**Provider terms (ask mode):** there is no auto-accept mode. When a turn ends on a terms offer, the status result carries `pendingApproval` (`kind: "terms"`) and `exchange.requiresAction` with the `approvalId` and provider requirements. Any `terms/accept` descriptor in that API payload is unavailable through MCP. To use the provider: 1. Show the user the terms (`terms/show` through `firecrawl_scrape` with `alexandria`). -2. Get the user's explicit consent to that provider's terms. A data request is not consent. -3. Run the `terms/accept` call through `firecrawl_scrape`. -4. Continue the same thread: call `firecrawl_agent` with the same `threadId` and `exchange.approve: { "approvalId": "..." }`. +2. Direct an organization admin to accept the terms at the provider's URL, or [data sources settings](https://www.firecrawl.dev/app/settings?tab=data-sources). A data request is not consent. +3. Wait for the admin to confirm acceptance in the dashboard. +4. Continue the same thread: call `firecrawl_agent` with the same `threadId` and `exchange.approve: { "approvalId": "..." }`. This resumes research and does not accept terms. If the user says no, call `firecrawl_agent` with the same `threadId` and `exchange.decline: { "approvalId": "..." }` instead. @@ -819,13 +819,13 @@ Then poll with `firecrawl_agent_status` using the returned job ID. } ``` -**Usage Example (continue the thread after the user accepted a provider's terms):** +**Usage Example (continue the thread after an admin confirmed dashboard acceptance):** ```json { "name": "firecrawl_agent", "arguments": { - "prompt": "I accepted the Apollo terms. Continue.", + "prompt": "The admin confirmed acceptance of the Apollo terms in the dashboard. Continue.", "threadId": "0199a1b2-0000-7000-8000-000000000031", "exchange": { "approve": { "approvalId": "0199a1b2-0000-7000-8000-000000000033" } } } diff --git a/src/index.ts b/src/index.ts index 7f432c72..10c44adb 100644 --- a/src/index.ts +++ b/src/index.ts @@ -3896,7 +3896,7 @@ const agentExchangeSchema = z }) .optional() .describe( - 'Answer yes to the pendingApproval the previous turn of this thread ended on (its id, also exchange.requiresAction.approvalId). Needs threadId. For a terms offer, send it only after the user explicitly agreed and terms/accept succeeded; callIds and always are ignored on terms offers. For paid calls, callIds picks a subset (default all) and always stops asking for the rest of the thread.' + 'Answer yes to the pendingApproval the previous turn of this thread ended on (its id, also exchange.requiresAction.approvalId). Needs threadId. For a terms offer, send it only after an organization admin confirms acceptance in the Firecrawl dashboard; this does not accept terms. callIds and always are ignored on terms offers. For paid calls, callIds picks a subset (default all) and always stops asking for the rest of the thread.' ), decline: z .strictObject({ approvalId: z.string().uuid() }) @@ -3908,7 +3908,7 @@ const agentExchangeSchema = z .enum(['skip', 'ask']) .optional() .describe( - 'What to do when a provider the agent would use needs data terms the team has not accepted. Gated providers are never called. "skip" (default): answer with accepted providers and list the rest in exchange.skippedProviders. "ask": the same, plus a terms pendingApproval and exchange.requiresAction with the terms/show and terms/accept calls. Each provider digest is string | null and always present; when null, terms/show returns it. There is no auto-accept. Omitted on a follow-up keeps the previous turn\'s value.' + 'What to do when a provider the agent would use needs data terms the team has not accepted. Gated providers are never called. "skip" (default): answer with accepted providers and list the rest in exchange.skippedProviders. "ask": the same, plus a terms pendingApproval and exchange.requiresAction. Read terms with terms/show; an organization admin accepts them in the Firecrawl dashboard. There is no auto-accept. Omitted on a follow-up keeps the previous turn\'s value.' ), }) .describe( @@ -3924,13 +3924,13 @@ server.addTool({ destructiveHint: false, // Gathers information only; does not delete external data or user resources. }, description: ` -Run web research that returns structured data when the URLs are not known or the answer spans several sites. Describe the fields you need in \`prompt\`, optionally pass a JSON \`schema\` and seed \`urls\`, and the research agent searches, navigates, reads pages, and returns JSON assembled across sources. Use it to research an entity plus its fields (founders, pricing, contact details), to build lists and datasets (companies, people, products, jobs, papers), and for pages that need navigation or interaction to reach the data. Optional \`effort\` sets the reasoning budget, \`maxCredits\` caps spend, and \`strictConstrainToURLs\` keeps the agent to the supplied \`urls\`. +Run web research when the URLs are unknown or the answer spans several sites. Describe the fields in \`prompt\`, optionally pass a JSON \`schema\` and seed \`urls\`; the agent searches, navigates, reads pages, and returns structured data across sources. -This call returns only a job ID, not the research result. Read the job with \`firecrawl_agent_status\` until it reaches \`completed\` or \`failed\`; a typical research run takes one to three minutes. For one known URL use \`firecrawl_scrape\` (with formats: ["json"] for structured output); for a plain lookup that a results page answers, use \`firecrawl_search\`. +This call returns only a job ID, not the research result. Poll \`firecrawl_agent_status\` until \`completed\` or \`failed\`; a typical run takes one to three minutes. For one known URL use \`firecrawl_scrape\`; for a lookup answered by a results page, use \`firecrawl_search\`. The job also returns a \`threadId\`. To continue that thread, pass it with a follow-up \`prompt\`; omitted \`mode\`, \`urls\`, \`schema\` and exchange settings carry over from the previous turn. -The agent only calls Alexandria providers whose data terms the team has accepted. The status result's \`exchange.skippedProviders\` lists gated providers that would have helped. With \`exchange.onTermsRequired\` "ask", a terms offer ends the turn: \`pendingApproval\` (kind "terms") and \`exchange.requiresAction\` carry the \`approvalId\` and the exact terms/show and terms/accept calls. Show the user the terms, get their EXPLICIT consent, run terms/accept through \`firecrawl_scrape\`, then call \`firecrawl_agent\` with the same \`threadId\` and \`exchange.approve: {approvalId}\`. If they decline, send \`exchange.decline: {approvalId}\` instead. Never call terms/accept without that consent; a data request is not consent. +The agent only calls Alexandria providers whose data terms the team has accepted. The status result's \`exchange.skippedProviders\` lists gated providers that would have helped. With \`exchange.onTermsRequired\` "ask", a terms offer ends the turn: \`pendingApproval\` (kind "terms") and \`exchange.requiresAction\` carry the \`approvalId\` and provider requirements. Read the terms with terms/show through \`firecrawl_scrape\` and present them to the user. Terms are accepted in the Firecrawl dashboard, not through this connection: an organization admin must accept them at the provider's URL or ${DATA_SOURCES_SETTINGS_URL}. Ignore any terms/accept call in the API response. Only after the admin confirms acceptance, call \`firecrawl_agent\` with the same \`threadId\` and \`exchange.approve: {approvalId}\` to resume; this does not accept terms. If they decline, send \`exchange.decline: {approvalId}\` instead. Never infer acceptance from a data request. `, outputSchema: agentOutputSchema, parameters: z.object({ diff --git a/src/tool-output.ts b/src/tool-output.ts index 3222882f..b035020e 100644 --- a/src/tool-output.ts +++ b/src/tool-output.ts @@ -248,10 +248,10 @@ export const agentStatusOutputSchema = z message: unknown('The agent\'s reply; in chat mode, the short answer to a follow-up.'), suggestions: unknown('Follow-ups the agent offers; send one as the prompt of the next turn with this threadId.'), exchange: unknown( - 'What the job did with Alexandria providers: onTermsRequired, paidCalls, creditsUsed, skippedProviders (gated providers that would have helped), and requiresAction (approvalId plus the terms/show and terms/accept calls, each provider digest string | null and always present; call accept only with the user\'s explicit consent, then continue the thread with exchange.approve: {approvalId}).' + 'What the job did with Alexandria providers: onTermsRequired, paidCalls, creditsUsed, skippedProviders (gated providers that would have helped), and requiresAction (approvalId and provider requirements). Read terms with terms/show. Ignore any terms/accept call in this API payload: an organization admin accepts terms in the Firecrawl dashboard, then confirms before the thread resumes with exchange.approve: {approvalId}.' ), pendingApproval: unknown( - 'Set when the job ended waiting on the caller. Answer it by calling firecrawl_agent with this threadId and exchange.approve or exchange.decline carrying its id. kind "terms" lists providers whose data terms need accepting; otherwise calls lists paid calls waiting for approval.' + 'Set when the job ended waiting on the caller. Answer it by calling firecrawl_agent with this threadId and exchange.approve or exchange.decline carrying its id. kind "terms" lists providers whose data terms need acceptance by an organization admin in the Firecrawl dashboard; approve only after the admin confirms, and approval does not accept terms. Otherwise calls lists paid calls waiting for approval.' ), }) .describe('Progress or final result of a research agent job.'); diff --git a/tests/mcp-smoke.test.mjs b/tests/mcp-smoke.test.mjs index a663ea51..518e7fc7 100644 --- a/tests/mcp-smoke.test.mjs +++ b/tests/mcp-smoke.test.mjs @@ -4445,7 +4445,15 @@ test('firecrawl_agent forwards onTermsRequired and status keeps the terms-requir assert.deepEqual(agentTool.inputSchema.properties.exchange.properties.onTermsRequired.enum, ['skip', 'ask']); assert.match(agentTool.description, /exchange\.skippedProviders/); assert.match(agentTool.description, /exchange\.requiresAction/); - assert.match(agentTool.description, /get their EXPLICIT consent.*Never call terms\/accept without that consent/); + assert.match(agentTool.description, /organization admin must accept.*app\/settings\?tab=data-sources/); + assert.match(agentTool.description, /Only after the admin confirms acceptance/); + assert.match(agentTool.description, /Ignore any terms\/accept call in the API response/); + assert.doesNotMatch(agentTool.description, /run terms\/accept through/); + assert.match(agentTool.inputSchema.properties.exchange.properties.approve.description, /this does not accept terms/); + assert.match(agentTool.inputSchema.properties.exchange.properties.onTermsRequired.description, /organization admin accepts them in the Firecrawl dashboard/); + const statusTool = tools.find((tool) => tool.name === 'firecrawl_agent_status'); + assert.match(statusTool.outputSchema.properties.exchange.description, /Ignore any terms\/accept call/); + assert.match(statusTool.outputSchema.properties.pendingApproval.description, /approval does not accept terms/); const asked = await client.request('tools/call', { arguments: { prompt: 'Find the key business contact at exa.ai', exchange: { onTermsRequired: 'ask' } }, @@ -4534,8 +4542,8 @@ test('firecrawl_agent answers a pending approval on a thread', async (t) => { assert.ok(agentTool.description.length <= CLAUDE_CODE_TEXT_CAP, `description is ${agentTool.description.length} chars`); assert.match(agentTool.description, /same `threadId` and `exchange\.approve: \{approvalId\}`/); assert.match(agentTool.description, /`exchange\.decline: \{approvalId\}`/); - assert.match(agentTool.description, /EXPLICIT consent/); - assert.match(agentTool.description, /Never call terms\/accept without that consent/); + assert.match(agentTool.description, /Terms are accepted in the Firecrawl dashboard/); + assert.match(agentTool.description, /Never infer acceptance from a data request/); const call = (args) => client.request('tools/call', { arguments: args, name: 'firecrawl_agent' }); @@ -4546,10 +4554,10 @@ test('firecrawl_agent answers a pending approval on a thread', async (t) => { assert.equal(followUp.structuredContent.threadTurn, 1); // 2. Answering a pending approval on the thread: forwarding of approve and - // decline only. terms/accept itself goes through firecrawl_scrape and is - // covered by the Alexandria terms tests; this does not run it. + // decline only. Terms acceptance happens in the dashboard, and the + // Alexandria terms tests cover its rejection through firecrawl_scrape. const approved = await call({ - prompt: 'I accepted the Apollo terms. Continue.', + prompt: 'The admin confirmed acceptance of the Apollo terms in the dashboard. Continue.', threadId, exchange: { approve: { approvalId } }, }); @@ -4594,7 +4602,7 @@ test('firecrawl_agent answers a pending approval on a thread', async (t) => { }); assert.deepEqual(bodies, [ { prompt: 'Only keep the founders', threadId, mode: 'chat' }, - { prompt: 'I accepted the Apollo terms. Continue.', threadId, exchange: { approve: { approvalId } } }, + { prompt: 'The admin confirmed acceptance of the Apollo terms in the dashboard. Continue.', threadId, exchange: { approve: { approvalId } } }, { prompt: 'Do not use Apollo.', threadId, exchange: { decline: { approvalId } } }, { prompt: 'Run only the first call.', From fc6cce399686de090f7b12469a9be5c7a3e49dea Mon Sep 17 00:00:00 2001 From: Max Loffgren Date: Wed, 30 Sep 2026 22:23:39 -0500 Subject: [PATCH 3/3] docs(mcp): clarify read-only browser profile parameters --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 02ab5f6f..a0d3e615 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ - Server instructions are one short string per surface (full, search, keyless). Each stays under 1,024 characters, routes search and scrape within its first 512, and names only tools that session lists. Hosted `/v2/mcp` now selects them per session, so a session with an API key or OAuth token gets the full-surface instructions instead of the keyless ones. Locally, only stdio without an API key, OAuth token, or `FIRECRAWL_API_URL` gets the keyless instructions; a self-hosted `FIRECRAWL_API_URL` and the local HTTP transport, which requires credentials or `FIRECRAWL_API_URL`, get the full-surface instructions. - An organization admin now accepts Alexandria provider terms in the Firecrawl dashboard. `firecrawl_scrape` refuses `terms/accept` and every other `terms/*` capability except `terms/show`, and terms errors link to `requiresAction.url` or the data sources settings page. Reading terms with `terms/show` is unchanged. -- On the hosted server, `firecrawl_scrape` is annotated `readOnlyHint: true` again. There, `firecrawl_scrape` and `firecrawl_search` `scrapeOptions` load a named `profile` without saving changes to it; save browser state with `firecrawl_interact` and `scrapeOptions.profile`. Local scrape and search keep browser actions and writable profiles and are annotated `readOnlyHint: false`. +- On the hosted server, `firecrawl_scrape` is annotated `readOnlyHint: true` again. There, `firecrawl_scrape` uses the top-level `profile` parameter and `firecrawl_search` uses `scrapeOptions.profile` to load saved browser state without saving changes to it; save browser state with `firecrawl_interact` and `scrapeOptions.profile`. Local scrape and search keep browser actions and writable profiles and are annotated `readOnlyHint: false`. - `tools/list` now sends each tool's top-level `title` (MCP 2025-06-18), copied from `annotations.title`. The title wording is unchanged. Clients that read only the top-level field, such as Codex's tool search, now see the same display names. - The npm package now bundles the pnpm-patched fastmcp. npm does not apply pnpm patches, so `npx firecrawl-mcp` installs were loading unpatched fastmcp from the registry, without the top-level titles or the `canList` and `beforeValidate` hooks. fastmcp's runtime imports (`@modelcontextprotocol/sdk`, `fuse.js`, `hono`, `mcp-proxy`, `undici`, `uri-templates`, `xsschema`) are now direct dependencies so they resolve under pnpm's non-hoisted layout too. - Keyless recovery messages now link to the caller's own signup link, `https://firecrawl.dev/k/` (a 12-character encrypted token), instead of `/app/api-keys`. The API issues the link (the `signup_url` of a keyless 429, or `signupUrl` from the eligibility check, which the server now also asks for when a keyless session calls an account-only tool) and the site decrypts it to MCP keyless attribution. When the API can't give one it sends the regular keyless signup link, which is relayed as is; with no API link at all the message uses the regular MCP signup link (`signin?utm_source=keyless&utm_medium=mcp`). Recovery payloads carry the link as `signup_url`. Signed-in users who open it land on the API keys page.