diff --git a/README.md b/README.md index 2e1f0bdc..2836f83d 100644 --- a/README.md +++ b/README.md @@ -37,7 +37,7 @@ A Model Context Protocol (MCP) server that brings [Firecrawl](https://github.com - Use `firecrawl_credit_usage` to check credits left or monthly consumption, optionally broken down by API key. - Consider something else when you need to hold a browser session open across many of your own steps with your own retry and termination logic: each `firecrawl_interact` call runs one `prompt` or `code` turn to completion and returns control — the session can persist across calls via `scrapeId` and ends with `firecrawl_interact_stop`, but you cannot drive it interactively step-by-step from the client side within a single call. -This server lists 26 tools when the full profile registers with default settings (feedback tools included, not running in local-keyless mode). Setting `FIRECRAWL_NO_SEARCH_FEEDBACK=1` and/or `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` removes the corresponding feedback tools and reduces this count, as does local keyless startup. For clients with a tool-slot limit: the hosted keyless endpoint (`https://mcp.firecrawl.dev/v2/mcp`, no API key) exposes only 3 — `firecrawl_scrape`, `firecrawl_search`, `firecrawl_parse` — and the dedicated [search-only endpoint](#search-only-endpoint) (`https://mcp.firecrawl.dev/v2/mcp-search`) exposes a fixed set of 8 tools (search, developer and research search, plus Alexandria catalogue lookup and execution). +Authenticated sessions expose the full tool set. Setting `FIRECRAWL_NO_SEARCH_FEEDBACK=1` and/or `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` hides the corresponding authenticated feedback tools. For clients with a tool-slot limit: the hosted keyless endpoint (`https://mcp.firecrawl.dev/v2/mcp`, no API key) exposes 4 tools: `firecrawl_scrape`, `firecrawl_search`, `firecrawl_parse`, and `firecrawl_feedback`. The dedicated [search-only endpoint](#search-only-endpoint) (`https://mcp.firecrawl.dev/v2/mcp-search`) exposes a fixed set of 8 tools (search, developer and research search, plus Alexandria catalogue lookup and execution). ## Installation @@ -508,7 +508,7 @@ For scientific papers, see [Research Tools](#12-research-tools-firecrawl_researc **Returns:** -- Array of search results (with optional scraped content), plus an `id` field. Pass that `id` to `firecrawl_search_feedback` after you've used the results to refund 1 credit (search costs 2) and improve search quality. +- Array of search results with optional scraped content, plus an `id` field. Keyless callers can use the returned job reference and feedback invitation with `firecrawl_feedback`. Authenticated callers can continue using `firecrawl_search_feedback` with its existing fields and policy. **Prompt Example:** @@ -558,17 +558,37 @@ Sends structured feedback on a previous `firecrawl_search` result. The first fee ### 3c. Generic Feedback Tool (`firecrawl_feedback`) -Sends structured feedback for a completed v2 endpoint job through `/v2/feedback`. -Use this for endpoint-level feedback on `scrape`, `parse`, `map`, or `search` -jobs. For search-result quality specifically, prefer -`firecrawl_search_feedback` because it includes search-specific guidance. +Sends evidence through `/v2/feedback`. Feedback on keyless Search, Scrape, and +Parse jobs is optional. Keyless guidance asks agents to submit concise feedback on observed result quality or missing coverage when the host permits it, especially if a result is wrong, incomplete, blocked, or an error. Feedback does not determine whether a task is complete. These jobs require `endpoint`, `jobId`, `rating`, `task`, `assessment`, and 1-20 +`observations`. Each observation has `kind`, `detail`, and `basis`: `output`, +`source_comparison`, or `expectation`. Source comparisons also require +`comparison: {reference, detail}`. + +The tool's `observations` parameter lists the category fields below. Use only available evidence +and keep unverified expectations distinct from source comparisons. + +Submit before the invitation's `expiresAt` deadline, which provides a 24-hour +feedback window for the job. Each job accepts one submission; +retrying it returns the original feedback ID. Feedback remains available after +operation allowance is exhausted and does not consume or restore that allowance. + +Authenticated callers retain the existing issue/note fields for Search, Scrape, +Parse, and Map. For authenticated Search-specific feedback, continue using +`firecrawl_search_feedback`. Choose the contract that matches the originating +job's authentication; adding credentials does not convert a keyless job. Keep feedback concise: use issue codes, tags, short notes, URLs, page numbers, and small metadata objects. Do not include raw scrape/parse outputs. -**Opt out:** set `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` (or `FIRECRAWL_DISABLE_ENDPOINT_FEEDBACK=1`) in the environment when starting the MCP server. The `firecrawl_feedback` tool will not be registered, so agents cannot call it. +Search observations identify delivered result positions or missing information. Scrape and Parse observations describe the requested output formats. Failed jobs use a `failure` observation based on the returned error. Parse additionally requires `docClass`: `born_digital`, `scanned`, `mixed`, or `unknown`. -**Usage Example:** +Task, assessment, and observation detail each require 10-2000 characters after trimming whitespace. Stored keyless feedback must fit within 8 KiB, including server defaults and verification flags. Submit from the same caller IP; attempts are rate limited. + +The tool's `observations` parameter lists the endpoint-specific categories, fields, and reason codes. See the [API feedback contract](https://docs.firecrawl.dev/api-reference/endpoint/feedback) for examples, format constraints, and Parse retention behavior. + +**Authenticated feedback preference:** set `FIRECRAWL_NO_ENDPOINT_FEEDBACK=1` (or `FIRECRAWL_DISABLE_ENDPOINT_FEEDBACK=1`) to hide `firecrawl_feedback` from authenticated sessions. Keyless sessions retain the tool and server-issued invitations regardless of these flags. The API includes a pointer on every eligible keyless job response. Feedback is optional; continued keyless access does not depend on it. + +**Authenticated usage example:** ```json { @@ -658,11 +678,11 @@ Check the status and results of an existing crawl job by ID. Parse local files or hosted upload references with Firecrawl's `/v2/parse` endpoint. -**Best for:** PDFs, Word documents, spreadsheets, HTML files, and other documents that need markdown or structured JSON output. Hosted MCP supports a two-step upload-ref flow; local direct file reads require a self-hosted `FIRECRAWL_API_URL`. +**Best for:** PDFs, Word documents, spreadsheets, HTML files, and other documents that need markdown or structured JSON output. Hosted MCP supports a two-step upload-ref flow; local MCP requires an explicit API URL before reading and uploading the requested file. **Not recommended for:** Remote URLs (use scrape), multiple files in one call (call parse once per file), or browser-only actions such as screenshots and clicks. -**Hosted MCP flow:** Hosted MCP cannot read the caller's filesystem directly. Call `firecrawl_parse` with `filePath` to receive a short-lived upload command and `nextToolCall`, upload the file locally, then call `firecrawl_parse` again with the returned `uploadRef`. Minting the hosted upload URL requires Firecrawl auth or keyless eligibility. In local `npx firecrawl-mcp` mode, direct file parsing currently requires `FIRECRAWL_API_URL` pointing to a self-hosted Firecrawl API; a plain cloud API-key-only local server cannot read and upload files through this tool. +**Hosted MCP flow:** Hosted MCP cannot read the caller's filesystem directly. Call `firecrawl_parse` with `filePath` to receive a short-lived upload command and `nextToolCall`, upload the file locally, then call `firecrawl_parse` again with the returned `uploadRef`. Minting the hosted upload URL requires Firecrawl auth or keyless eligibility. In local `npx firecrawl-mcp` mode, `FIRECRAWL_API_URL` must be explicitly configured before the server reads or uploads `filePath`. There is no default upload destination for local Parse. Both authenticated and eligible keyless calls are supported by the selected API. Running MCP locally does not perform parsing on the local machine. This configuration requirement selects the destination; it does not restrict which files the process can read. **Usage Example:** diff --git a/src/index.ts b/src/index.ts index 04302253..bf838175 100644 --- a/src/index.ts +++ b/src/index.ts @@ -782,13 +782,13 @@ async function authenticateRequest( !process.env.FIRECRAWL_API_KEY && !process.env.FIRECRAWL_API_URL ) { - // No credential and no self-hosted URL: run in keyless mode. scrape and - // search work for free (rate-limited per IP) against the Firecrawl cloud; - // every other tool needs an API key and will return Unauthorized. + // Without credentials or a custom API URL, use the cloud keyless tools. console.error( - 'No FIRECRAWL_API_KEY or FIRECRAWL_API_URL set — running in keyless mode. ' + - 'firecrawl_scrape and firecrawl_search are free (rate-limited per IP) against the Firecrawl cloud; ' + - 'other tools require an API key (get one free at https://firecrawl.dev).' + 'No FIRECRAWL_API_KEY or FIRECRAWL_API_URL set. Running in keyless mode. ' + + 'firecrawl_scrape and firecrawl_search use the Firecrawl cloud with usage limits. ' + + 'Local firecrawl_parse requires FIRECRAWL_API_URL to be explicitly configured before reading or uploading files. ' + + 'Consider submitting feedback through firecrawl_feedback, especially if a keyless result is wrong, incomplete, blocked, or an error. Include specific evidence to help improve Firecrawl. ' + + 'Other tools require an API key (get one free at https://firecrawl.dev).' ); } @@ -1314,7 +1314,7 @@ const openAiAppsChallengeToken = normalizeHeader( const FULL_PROFILE_INSTRUCTIONS = `Firecrawl provides web search, page retrieval, site URL discovery, multi-page collection, structured page data, monitoring, and multi-source research that returns structured data. Match the requested operation to the tool boundary: firecrawl_scrape retrieves one supplied page and can return JSON matching a supplied schema, firecrawl_map enumerates URLs under a site without retrieving their content, and firecrawl_agent runs multi-source research and returns structured data when the URLs are not known or the answer spans several sites (an entity plus its fields, a list, a dataset); its result is read with firecrawl_agent_status. Authenticated firecrawl_search returns web results together with matching Alexandria providers in data.tools. ${ALEXANDRIA_CATALOGUE_SENTENCE} A matching provider can return the same fields across several entities, provenance, exact figures or timestamps, or a large set of records through firecrawl_scrape with its published contract. If web results already answer the question, use them. For the same fields across multiple pages, firecrawl_find_tools offers free provider discovery. Use firecrawl_find_tools to read a contract that was not returned in full or to browse the catalogue by category. ${ALEXANDRIA_SOURCES_OPT_OUT} If no provider fits, continue with web search or firecrawl_agent. For biomedical, life-science, clinical, or arXiv literature, the firecrawl_research_* tools search a paper index of abstracts and full text; firecrawl_search with categories: ["research"] is a website filter over ordinary web results and reaches different sources. For a programming question (code behaviour, a library or framework, an API contract, an error message, or a known bug), firecrawl_developer_search (or firecrawl_search with categories: ["developer"]) searches an index of public repositories, GitHub issues, merged pull requests, READMEs, and code documentation. firecrawl_search with sources: [{type: "alexandria"}] returns compact tool summaries in data.tools; toolDetail: "full" includes contracts, firecrawl_find_tools starts with categories, lists providers, then compact tools, and expands the selected full contract, and firecrawl_scrape with alexandria: [{provider, capability, options}] executes up to ten capabilities and returns their results. Alexandria access needs an API key on a team with it enabled. Provide only the required inputs and account for stated network or external side effects.`; -const KEYLESS_PROFILE_INSTRUCTIONS = `Hosted keyless sessions expose firecrawl_search, firecrawl_scrape, and firecrawl_parse with usage limits. firecrawl_search searches the web. For programming questions, firecrawl_search with categories: ["developer"] searches indexed public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation. For biomedical, life-science, clinical, or arXiv literature, firecrawl_search with categories: ["research"] filters ordinary web results to research-affiliated websites. firecrawl_scrape retrieves one supplied page and can return JSON matching a supplied schema. firecrawl_parse processes supported local files through its two-phase upload flow. An Authorization bearer API key can provide higher usage limits and expose additional tools, subject to plan, deployment, and team policy, including firecrawl_map for site URL discovery, firecrawl_agent and firecrawl_agent_status for multi-source research that returns structured data when the URLs are not known, firecrawl_research_* for paper-index and repository research, and firecrawl_find_tools as the progressive Alexandria catalogue lookup alongside the Alexandria options of firecrawl_search and firecrawl_scrape for catalogued data providers.`; +const KEYLESS_PROFILE_INSTRUCTIONS = `Keyless sessions can use firecrawl_search, firecrawl_scrape, firecrawl_parse, and firecrawl_feedback with usage limits. Submit concise feedback through firecrawl_feedback on observed result quality or missing coverage when a keyless Search, Scrape, or Parse result includes a feedback invitation and the host permits it, especially if the result is wrong, incomplete, blocked, or an error. Use its job reference and evidence already available, and omit sensitive information. Feedback does not determine whether the task is complete and does not consume operation quota. If it is unavailable, declined, or rejected, continue without retries. firecrawl_search searches the web. For programming questions, firecrawl_search with categories: ["developer"] searches indexed public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation. For biomedical, life-science, clinical, or arXiv literature, firecrawl_search with categories: ["research"] filters ordinary web results to research-affiliated websites. firecrawl_scrape retrieves one supplied page and can return JSON matching a supplied schema. Hosted firecrawl_parse processes supported local files through its two-phase upload flow. Local firecrawl_parse requires FIRECRAWL_API_URL to be explicitly configured before reading or uploading files. An Authorization bearer API key can provide higher usage limits and expose additional tools, subject to plan, deployment, and team policy, including firecrawl_map for site URL discovery, firecrawl_agent and firecrawl_agent_status for multi-source research that returns structured data when the URLs are not known, firecrawl_research_* for paper-index and repository research, and firecrawl_find_tools as the progressive Alexandria catalogue lookup alongside the Alexandria options of firecrawl_search and firecrawl_scrape for catalogued data providers.`; // The search surface exposes web/developer/research search plus the two Alexandria // tools (catalogue lookup and provider execution). Its instructions @@ -1457,6 +1457,7 @@ const KEYLESS_TOOL_NAMES = new Set([ 'firecrawl_scrape', 'firecrawl_search', 'firecrawl_parse', + 'firecrawl_feedback', ]); function isHostedKeylessSession(session?: SessionData): boolean { @@ -1468,7 +1469,7 @@ function isHostedKeylessSession(session?: SessionData): boolean { } // A stdio client without a cloud credential can use only the keyless tools. -// Do this at registration time so unsupported feedback tools are not advertised. +// Apply the keyless tool set at registration time for local sessions. function isLocalKeylessStartup(): boolean { return ( process.env.CLOUD_SERVICE !== 'true' && @@ -1597,6 +1598,13 @@ async function runWithCredentialRecovery( // different fault and keeps its own reconnect guidance; a keyless session // never sent an account credential at all. if (session?.authType !== 'api-key' || !isCoreCredentialRejection(error)) { + // Job failures already expose the full API envelope in text and extras. + if (!hasCredential(session) && error instanceof UserError) { + const metadata = error.extras?.metadata; + if (metadata && typeof metadata === 'object' && 'jobId' in metadata) { + throw error; + } + } const hints = readErrorAgentHints(error); if (hints) { const message = error instanceof Error ? error.message : String(error); @@ -2908,11 +2916,7 @@ For a programming question, add \`categories: ["developer"]\`; its hits return i assertExchangeCredential(session); if (isKeylessMode(session)) { const json = await keylessPost('/v2/search', searchBody, session); - // Search feedback requires an authenticated account. Do not expose its - // identifier to keyless clients, where it would invite an unusable call. - const keylessResponse = { ...(json ?? {}) }; - delete keylessResponse.id; - return structuredCompact(keylessResponse); + return structuredCompact(json); } // Call /v2/search through the SDK's HTTP layer (auth + retries) instead // of `client.search()` so we preserve the full response envelope. The @@ -3202,7 +3206,7 @@ async function keylessPost( body: JSON.stringify(body), }); const json: any = await response.json().catch(() => ({})); - if (!response.ok) { + if (!response.ok || json?.success === false) { if (isKeylessMode(session) && response.status === 429) { // The API normally supplies requests|credits. Preserve a structured, // non-specific recovery payload during a skewed or legacy deployment. @@ -3221,6 +3225,12 @@ async function keylessPost( if (hints) payload.agent_hints = hints; throw new UserError(String(payload.message), payload); } + if (json?.metadata?.jobId) { + throw new UserError( + JSON.stringify(json, null, 2), + json + ); + } throw new CoreHttpError( json?.error || `Firecrawl request failed (HTTP ${response.status})`, response.status, @@ -3507,10 +3517,59 @@ Eligibility is limited to successful searches within the feedback age window. Th if (ENDPOINT_FEEDBACK_DISABLED) { console.error( - '[firecrawl-mcp] Endpoint feedback tool disabled by FIRECRAWL_NO_ENDPOINT_FEEDBACK; firecrawl_feedback will not be registered.' + '[firecrawl-mcp] Authenticated endpoint feedback tool disabled by FIRECRAWL_NO_ENDPOINT_FEEDBACK. Keyless feedback remains available.' ); } +// Keyless observation rules live on the observations parameter so the tool +// description stays within the client's description window. +const KEYLESS_OBSERVATIONS_CONTRACT = `Search: useful and irrelevant require a one-based position within the delivered group. source names the response group the position refers to: web, images, or news. It is required for multi-source jobs. Omission defaults to web, so images-only and news-only jobs must explicitly name their source. The position must exist in that requested group. irrelevant requires reason: aggregator_over_official, off_topic, stale, wrong_content_type, snippet_misleading, or blocked_or_paywalled. vertical is required on missing and optional on useful/irrelevant: web_general, social, business, research, developer, news, government, finance, or other. missing may include topic (up to 200 characters). missing and irrelevant may include knownSources (up to 20 HTTP(S) URLs): where absent content lives or the source that should have ranked instead. Unmentioned results are unassessed; a full ranking is not required. Do not submit engine attribution. + +Scrape: kind correct, wrong_success, incomplete, or incorrect. wrong_success requires reason: blocked_shell, login_required, paywall, empty, wrong_page, stale, or wrong_locale. incomplete requires reason: partial_content, dynamic_content, pagination, main_content_stripped, or format_lost. incorrect requires reason: wrong, hallucinated, or missing_fields. correct has no reason. Optional location is up to 200 characters. No retryOutcome. hallucinated applies only to json, deterministicJson, summary, question, highlights, and changeTracking in json mode; missing_fields applies only to json and deterministicJson. For incomplete and incorrect, prefer source_comparison when the source is already available. + +Parse: docClass is required once per submission: born_digital, scanned, mixed, or unknown. Observation kind: correct, text_ocr, table, formula, chart_figure, reading_order, headers_footers, headings_formatting, completeness, images_dropped, or incorrect. text_ocr requires reason: misread_chars, garbled, or missing_text. table requires reason: structure, cells_glued, or digits. completeness requires reason: pages_missing, truncated_at_max_pages, or sections_dropped. incorrect requires reason: wrong, hallucinated, or missing_fields. Other kinds have no reason subtype. Optional page is a one-based positive integer. incorrect applies to json and summary outputs. For text_ocr and table, include the correct text or cell values in comparison.detail when already known. Parse feedback does not automatically retain the document, extracted output, page images, or layout blocks; submitted observations and corrections are retained. + +Scrape and Parse observations other than failure: format must be a format type the job requested. It is required for output and source_comparison observations when multiple formats were requested; optional for expectation observations and single-format jobs. All observations retain detail and basis; source_comparison requires comparison: {reference, detail}. comparison.detail contains the correct content from the inspected source. + +Failed Search, Scrape, or Parse jobs: use kind failure with reason timeout, transport_error, proxy_error, or other. Accepted only for a failed job. Include detail and basis; do not supply position, source, format, location, or page. Parse still requires docClass (unknown is allowed). + +If the saved Search response is unavailable, otherwise valid observations are accepted and stored with metadata.unverified: true because their positions could not be checked. Job ownership and requested sources are still checked. Available results must contain every referenced position. + +Reason definitions: +- aggregator_over_official: An intermediary was returned where the task needed an available official or primary source. +- off_topic: The result addresses a different topic from the task. +- stale: The content is outdated for the time or version the task requires. +- wrong_content_type: The destination has the wrong content type for the task, such as a discussion instead of a reference. +- snippet_misleading: The returned description misrepresents source content already inspected. +- blocked_or_paywalled: Access to the destination was observed to be blocked or require a subscription; do not infer this from its URL or snippet. +- blocked_shell: The successful response contains a bot challenge or access-blocking shell instead of the requested content. +- login_required: The successful response contains a login requirement instead of the requested content. +- paywall: The successful response contains a subscription barrier instead of the requested content. +- empty: The successful response contains no meaningful requested content. +- wrong_page: The successful response contains a different page or resource. +- wrong_locale: The response uses the wrong language or region for the task. +- partial_content: Only part of the expected content was returned, without a more specific known cause. +- dynamic_content: Content loaded by client-side rendering or interaction is missing. +- pagination: Expected content on additional pages is missing. +- main_content_stripped: Content filtering removed requested primary content. +- format_lost: Text is present, but meaningful structure such as headings, lists, or code formatting was lost. +- wrong: Returned facts or values conflict with the inspected source. +- hallucinated: The output asserts content unsupported by the inspected source. +- missing_fields: Requested fields are absent from the structured output. +- misread_chars: Characters were recognized incorrectly. +- garbled: Extracted text is corrupted or unreadable. +- missing_text: Visible source text was omitted. +- structure: Table rows, columns, or header relationships were reconstructed incorrectly. +- cells_glued: Distinct table cells were merged. +- digits: Numeric table values were recognized incorrectly. +- pages_missing: Source pages are absent from the output. +- truncated_at_max_pages: Extraction ended at the configured page limit; this does not by itself imply a parser error. +- sections_dropped: Sections within processed pages were omitted. +- timeout: The operation explicitly reported a timeout. +- transport_error: The operation explicitly reported a network, connection, or TLS failure. +- proxy_error: The operation explicitly reported a proxy failure. +- other: Another operation failure was reported; describe the returned error without guessing its cause.`; + /** Whether firecrawl_feedback is registered on this process, so Alexandria results only point at a tool that exists. */ function alexandriaFeedbackAvailable(session?: SessionData): boolean { // The search surface does not register firecrawl_feedback, so its results @@ -3522,21 +3581,29 @@ function alexandriaFeedbackAvailable(session?: SessionData): boolean { ); } -if (alexandriaFeedbackAvailable()) { +if ( + !ENDPOINT_FEEDBACK_DISABLED || + !resolveCredentialFromEnv() || + process.env.CLOUD_SERVICE === 'true' +) { server.addTool({ name: 'firecrawl_feedback', + canList: (session: SessionData) => + !ENDPOINT_FEEDBACK_DISABLED || !hasCredential(session), annotations: { title: 'Firecrawl feedback', - readOnlyHint: false, // POSTs structured feedback for a completed job to /v2/feedback. + readOnlyHint: false, // POSTs structured feedback to /v2/feedback. openWorldHint: true, // Feedback is tied to jobs that processed open-web URLs. destructiveHint: false, // Additive only; submits ratings and notes, does not delete jobs or external content. }, description: ` -Submit concise quality feedback for a completed search, scrape, parse, or map job. Provide the endpoint, job ID, rating, and relevant issue codes or small contextual fields; omit large page contents and raw outputs. +Submit job quality feedback through /v2/feedback. Authenticated jobs use existing issue/note fields. Omit full page contents and raw outputs. For an Alexandria session, set endpoint to \`alexandria\`, omit jobId, and provide requestedWebsite (url and requestedFunctionality), objective, rationale, and rating. objective is the underlying goal behind the session: what you or your user were ultimately trying to accomplish (for example, "shortlist federal IT contracts to bid on this quarter"), not only what was needed from this website. Optional providerFeedback and capabilityFeedback describe gaps or errors. Capability issues: new_capability_request (requires requestedFunctionality), missing_capability, insufficient_functionality, incorrect_result, execution_error, other. Alexandria feedback has no job-age deadline and no credit refund. -Returns submission status, feedback ID, and accounting fields. +Submit concise keyless Search, Scrape, or Parse feedback on observed result quality or missing coverage when the host permits it, especially if a result is wrong, incomplete, blocked, or an error. Feedback does not determine whether the task is complete. It requires task, assessment, rating, and 1-20 observations. Task, assessment, and each detail contain 10-2000 characters. Each observation includes kind, detail, and basis: output, source_comparison, or expectation. A source_comparison also requires comparison: {reference, detail}. The observations parameter describes each category's fields and reason codes. + +Use evidence already available; no extra investigation. The stored submission must fit within 8 KiB, including server defaults. Submit from the same caller IP before the invitation's expiresAt deadline (24-hour feedback window for the job). One submission per job; retries return the original feedback ID. Attempts are rate limited. Feedback remains available after operation quota exhaustion and does not consume or restore quota. Contract: https://docs.firecrawl.dev/api-reference/endpoint/feedback. `, outputSchema: feedbackOutputSchema, parameters: z.object({ @@ -3544,6 +3611,18 @@ Returns submission status, feedback ID, and accounting fields. jobId: z.string().uuid('jobId must be the UUID returned by Firecrawl').optional(), ...alexandriaFeedbackFields, rating: z.enum(['good', 'bad', 'partial']), + task: z.string().trim().min(10).max(2000).optional(), + assessment: z.string().trim().min(10).max(2000).optional(), + docClass: z + .enum(['born_digital', 'scanned', 'mixed', 'unknown']) + .optional() + .describe('Required once per keyless Parse submission.'), + observations: z + .array(z.record(z.string(), z.unknown())) + .min(1) + .max(20) + .optional() + .describe(KEYLESS_OBSERVATIONS_CONTRACT), issues: z.array(feedbackIssueSchema).max(20).optional(), tags: z.array(feedbackIssueSchema).max(20).optional(), note: z.string().max(4000).optional(), @@ -3573,10 +3652,19 @@ Returns submission status, feedback ID, and accounting fields. { session, log, client: mcpClient } ): Promise => { const origin = requestOrigin(mcpClient, session); + if (ENDPOINT_FEEDBACK_DISABLED && hasCredential(session)) { + throw new UserError( + 'Endpoint feedback is disabled for authenticated sessions.' + ); + } const { endpoint, jobId, rating, + task, + assessment, + docClass, + observations, issues, tags, note, @@ -3590,6 +3678,10 @@ Returns submission status, feedback ID, and accounting fields. endpoint: 'search' | 'scrape' | 'parse' | 'map' | 'alexandria'; jobId: string; rating: 'good' | 'bad' | 'partial'; + task?: string; + assessment?: string; + docClass?: 'born_digital' | 'scanned' | 'mixed' | 'unknown'; + observations?: Record[]; issues?: string[]; tags?: string[]; note?: string; @@ -3610,8 +3702,17 @@ Returns submission status, feedback ID, and accounting fields. const credential = credentialForOutboundRequest(session); if (credential) { headers['Authorization'] = `Bearer ${credential}`; - } else if (process.env.CLOUD_SERVICE === 'true') { - throw new Error('Unauthorized: missing API key for feedback.'); + } else if (isHostedKeylessSession(session)) { + if (!session?.keylessClientIp || !process.env.KEYLESS_PROXY_SECRET) { + return structuredText({ + success: false, + error: 'Feedback requires a trusted client identity.', + retryable: false, + }); + } + headers['x-firecrawl-keyless-ip'] = session.keylessClientIp; + headers['x-firecrawl-keyless-secret'] = + process.env.KEYLESS_PROXY_SECRET; } const body = endpoint === 'alexandria' @@ -3620,6 +3721,10 @@ Returns submission status, feedback ID, and accounting fields. endpoint, jobId, rating, + task, + assessment, + docClass, + observations, issues, tags, note, @@ -3632,6 +3737,26 @@ Returns submission status, feedback ID, and accounting fields. origin, }); + if (isKeylessMode(session)) { + if (!['search', 'scrape', 'parse'].includes(endpoint)) { + throw new UserError( + 'Keyless feedback supports Search, Scrape, and Parse. Other endpoints require authentication.' + ); + } + const required = { + task, + assessment, + observations, + ...(endpoint === 'parse' ? { docClass } : {}), + }; + const missing = Object.entries(required) + .filter(([, value]) => value === undefined) + .map(([name]) => name); + if (missing.length) { + throw new UserError(`Keyless feedback requires ${missing.join(', ')}.`); + } + } + log.info('Submitting endpoint feedback', { endpoint, jobId, rating }); const response = await fetch(`${apiBase}/v2/feedback`, { method: 'POST', @@ -3666,7 +3791,11 @@ Returns submission status, feedback ID, and accounting fields. status: response.status, feedbackErrorCode: parsed?.feedbackErrorCode, error: parsed?.error ?? `HTTP ${response.status}`, - retryable: response.status >= 500, + retryable: response.status >= 500 || (!credential && response.status === 429), + ...(!credential ? { + details: parsed?.details, + retry_after_seconds: parsed?.retry_after_seconds, + } : {}), ...(readAgentHints(parsed) ? { agent_hints: readAgentHints(parsed) } : {}), @@ -4226,7 +4355,7 @@ server.addTool({ description: ` Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with \`pdfOptions.maxPages\`. -Local MCP reads \`filePath\` from the server filesystem. Hosted MCP uses two calls: first provide \`filePath\` to receive upload instructions, upload locally, then call again with the returned \`uploadRef\`; do not send both fields together. Remote web URLs belong in \`firecrawl_scrape\`. +Local MCP requires an explicitly configured \`FIRECRAWL_API_URL\` before reading \`filePath\` from the server filesystem and uploading it. Parsing happens on that API server. Hosted MCP uses two calls: first provide \`filePath\` to receive upload instructions, upload locally, then call again with the returned \`uploadRef\`; do not send both fields together. Remote web URLs belong in \`firecrawl_scrape\`. Set \`redactPII\` to request redaction of personally identifiable information in the returned content. \`zeroDataRetention\` requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Authenticated final responses can include a \`data.metadata.scrapeId\` for optional parse feedback. `, @@ -4245,8 +4374,8 @@ Set \`redactPII\` to request redaction of personally identifiable information in const apiUrl = process.env.FIRECRAWL_API_URL; if (!apiUrl) { - throw new Error( - 'firecrawl_parse requires FIRECRAWL_API_URL to be set to a self-hosted Firecrawl API instance.' + throw new UserError( + 'Local firecrawl_parse requires FIRECRAWL_API_URL to be explicitly configured before reading or uploading files.' ); } @@ -4302,25 +4431,28 @@ Set \`redactPII\` to request redaction of personally identifiable information in }); const responseText = await response.text(); - if (!response.ok) { - let parsed: unknown; - try { - parsed = JSON.parse(responseText); - } catch { - // Keep the original non-JSON error text. + let result: any; + try { + result = JSON.parse(responseText); + } catch { + result = undefined; + } + if (!response.ok || (!hasCredential(session) && result?.success === false)) { + if (!hasCredential(session) && result?.metadata?.jobId) { + throw new UserError( + JSON.stringify(result, null, 2), + result + ); } throw new CoreHttpError( `Parse request failed with status ${response.status}: ${responseText}`, response.status, - readAgentHints(parsed) + readAgentHints(result) ); } - - try { - return structuredText(JSON.parse(responseText)); - } catch { - return withStructured(responseText, { raw: responseText }); - } + return result === undefined + ? withStructured(responseText, { raw: responseText }) + : structuredText(result); }, }); diff --git a/src/tool-output.ts b/src/tool-output.ts index 57f83fc5..74f6b2ce 100644 --- a/src/tool-output.ts +++ b/src/tool-output.ts @@ -160,7 +160,8 @@ export const searchOutputSchema = z data: unknown('Ranked results grouped by source, such as `web`, `news`, `images`, and `alexandria`.'), error, warning, - id: str('Search identifier, for optional `firecrawl_search_feedback`.'), + id: str('Search job identifier: keyless callers use `firecrawl_feedback`; authenticated callers can use `firecrawl_search_feedback`.'), + metadata: unknown('Keyless job reference and optional feedback invitation.'), creditsUsed: num('Credits this search consumed.'), tools: unknown('Domain-matched Alexandria tools for the results.'), nextTool: unknown('A follow-up tool call that continues this search.'), @@ -185,6 +186,8 @@ export const feedbackOutputSchema = z error, status: num('HTTP status when the submission was rejected.'), feedbackErrorCode: str('Machine-readable reason the submission was rejected.'), + details: unknown('Keyless request validation errors.'), + retry_after_seconds: num('Seconds to wait before retrying a throttled keyless submission.'), retryable: bool('Whether retrying the submission can succeed.'), message: str('Human-readable result of the submission.'), feedbackId: str('Identifier of the recorded feedback.'), diff --git a/tests/mcp-smoke.test.mjs b/tests/mcp-smoke.test.mjs index 47afff78..f68bbf43 100644 --- a/tests/mcp-smoke.test.mjs +++ b/tests/mcp-smoke.test.mjs @@ -3,6 +3,9 @@ import { spawn } from 'node:child_process'; import { readFileSync } from 'node:fs'; import { createServer } from 'node:http'; import net from 'node:net'; +import { mkdtemp, writeFile, rm } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import { join, relative } from 'node:path'; import test from 'node:test'; import { setTimeout as delay } from 'node:timers/promises'; import { assertAgentMetadataPolicy } from '../scripts/agent-metadata-policy.mjs'; @@ -676,6 +679,9 @@ async function startFakeFirecrawlBackend(options = {}) { keylessEligible = false, keylessEligibilityResponse, searchResponse, + scrapeResponse, + feedbackResponse, + parseResponse, } = options; const requests = []; const server = createServer(async (req, res) => { @@ -751,6 +757,20 @@ async function startFakeFirecrawlBackend(options = {}) { return; } + if (req.method === 'POST' && req.url === '/v2/feedback') { + const response = feedbackResponse ?? { + status: 200, + body: { + success: true, + feedbackId: '00000000-0000-4000-8000-000000000102', + creditsRefunded: 0, + }, + }; + res.writeHead(response.status, { 'content-type': 'application/json' }); + res.end(JSON.stringify(response.body)); + return; + } + if (req.method === 'POST' && req.url === '/v2/search') { if (searchResponse) { res.writeHead(searchResponse.status, { 'content-type': 'application/json' }); @@ -769,6 +789,12 @@ async function startFakeFirecrawlBackend(options = {}) { return; } + if (req.method === 'POST' && req.url === '/v2/scrape' && scrapeResponse) { + res.writeHead(scrapeResponse.status, { 'content-type': 'application/json' }); + res.end(JSON.stringify(scrapeResponse.body)); + return; + } + if (req.method === 'GET' && req.url === '/v2/team/credit-usage') { res.writeHead(200, { 'content-type': 'application/json' }); res.end( @@ -825,6 +851,11 @@ async function startFakeFirecrawlBackend(options = {}) { } if (req.method === 'POST' && req.url === '/v2/parse') { + if (parseResponse) { + res.writeHead(parseResponse.status, { 'content-type': 'application/json' }); + res.end(JSON.stringify(parseResponse.body)); + return; + } res.writeHead(200, { 'content-type': 'application/json' }); res.end( JSON.stringify({ @@ -840,18 +871,6 @@ async function startFakeFirecrawlBackend(options = {}) { return; } - if (req.method === 'POST' && req.url === '/v2/feedback') { - res.writeHead(200, { 'content-type': 'application/json' }); - res.end( - JSON.stringify({ - creditsRefunded: 0, - feedbackId: '00000000-0000-4000-8000-000000000102', - success: true, - }) - ); - return; - } - if (req.method === 'GET' && req.url?.startsWith('/v2/monitor')) { res.writeHead(200, { 'content-type': 'application/json' }); res.end(JSON.stringify({ success: true, data: [] })); @@ -942,7 +961,7 @@ test('HTTP cloud keyless transport preserves app challenge without advertising O const anonymousTools = parseSseJson(await unauthenticated.text()).result.tools; assert.deepEqual( anonymousTools.map((tool) => tool.name).sort(), - ['firecrawl_parse', 'firecrawl_scrape', 'firecrawl_search'] + ['firecrawl_feedback', 'firecrawl_parse', 'firecrawl_scrape', 'firecrawl_search'] ); const anonymousParse = anonymousTools.find( (tool) => tool.name === 'firecrawl_parse' @@ -1567,8 +1586,9 @@ test('credit usage tool exposes current balance and both historical request shap } }); -test('local keyless stdio keeps profile guidance keyless-scoped and omits feedback tools', async (t) => { +test('local keyless stdio keeps profile guidance keyless-scoped and exposes shared feedback', async (t) => { const child = spawnServer({ + FIRECRAWL_NO_ENDPOINT_FEEDBACK: '1', FIRECRAWL_API_KEY: '', FIRECRAWL_API_URL: '', FIRECRAWL_OAUTH_TOKEN: '', @@ -1588,7 +1608,7 @@ test('local keyless stdio keeps profile guidance keyless-scoped and omits feedba const apiKeyGuidance = init.instructions.slice(apiKeyBoundaryIndex); assert.match( keylessGuidance, - /Hosted keyless sessions expose firecrawl_search, firecrawl_scrape, and firecrawl_parse with usage limits/i + /Keyless sessions can use firecrawl_search, firecrawl_scrape, firecrawl_parse, and firecrawl_feedback with usage limits/i ); assert.match( keylessGuidance, @@ -1599,7 +1619,10 @@ test('local keyless stdio keeps profile guidance keyless-scoped and omits feedba /firecrawl_search with categories: \["research"\].*research-affiliated websites/i ); assert.match(keylessGuidance, /firecrawl_scrape retrieves one supplied page/i); - assert.match(keylessGuidance, /firecrawl_parse processes supported local files/i); + assert.match( + keylessGuidance, + /Hosted firecrawl_parse.*two-phase upload flow.*Local firecrawl_parse requires FIRECRAWL_API_URL.*before reading or uploading files/i + ); assert.doesNotMatch( keylessGuidance, /firecrawl_(?:map|agent|agent_status|research_)/i @@ -1621,12 +1644,18 @@ test('local keyless stdio keeps profile guidance keyless-scoped and omits feedba const tools = await client.request('tools/list'); const toolNames = tools.tools.map((tool) => tool.name); assert.equal(toolNames.includes('firecrawl_search_feedback'), false); - assert.equal(toolNames.includes('firecrawl_feedback'), false); + assert.equal(toolNames.includes('firecrawl_feedback'), true); const search = tools.tools.find((tool) => tool.name === 'firecrawl_search'); assert.ok(search); + // Tool descriptions are shared by every session; keyless feedback is + // pointed to in keyless results, the keyless instructions and the feedback + // tool instead. + for (const name of ['firecrawl_search', 'firecrawl_scrape', 'firecrawl_parse']) { + assert.doesNotMatch(tools.tools.find(tool => tool.name === name).description, /firecrawl_feedback/); + } assert.match( - search.description, - /authenticated responses can include an `id` for optional search feedback/i + keylessGuidance, + /Submit concise feedback through firecrawl_feedback.*Feedback does not determine whether the task is complete/i ); }); @@ -2347,7 +2376,7 @@ test('HTTP cloud transport serves an eligible keyless client and forwards its IP const message = parseSseJson(await toolCall.text()); assert.notEqual(message.result.isError, true); const keylessSearchPayload = JSON.parse(message.result.content[0].text); - assert.equal('id' in keylessSearchPayload, false); + assert.equal(keylessSearchPayload.id, '00000000-0000-4000-8000-000000000000'); const eligibilityCalls = backend.requests.filter( (r) => r.url === '/v2/keyless/eligibility' @@ -4337,6 +4366,420 @@ test('account OAuth tokens cannot replay on keyless and invalid keys get correct assert.deepEqual(introspectedTokens, ['fco_account']); }); + +test('hosted keyless feedback bypasses exhausted operation allowance and preserves rejection details', async (t) => { + for (const status of [200, 400, 429, 503]) { + await t.test(`HTTP ${status}`, async (t) => { + const body = + status === 200 + ? { + success: true, + feedbackId: '00000000-0000-4000-8000-000000000001', + creditsRefunded: 0, + } + : status === 429 + ? { + success: false, + error: 'Too many feedback attempts. Retry in one minute.', + retry_after_seconds: 60, + } + : { + success: false, + error: 'Feedback rejected', + feedbackErrorCode: 'INVALID_BODY', + }; + const backend = await startFakeFirecrawlBackend({ + keylessEligible: false, + feedbackResponse: { status, body }, + }); + t.after(() => backend.close()); + const port = await getFreePort(); + const child = spawnServer({ + FIRECRAWL_DISABLE_ENDPOINT_FEEDBACK: '1', + CLOUD_SERVICE: 'true', + FASTMCP_ENDPOINT: '/v2/mcp', + FIRECRAWL_API_URL: backend.url, + FIRECRAWL_OAUTH_ISSUER: backend.url, + HTTP_STREAMABLE_SERVER: 'true', + PORT: String(port), + KEYLESS_PROXY_SECRET: 'feedback-test-secret', + }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + const response = await httpToolCall(port, { + id: `feedback-${status}`, + headers: { 'x-forwarded-for': '203.0.113.71' }, + params: { + name: 'firecrawl_feedback', + arguments: { + endpoint: 'search', + jobId: '00000000-0000-4000-8000-000000000000', + rating: 'partial', + task: 'Read the API retry reference', + assessment: 'The reference explains supported retry intervals.', + observations: [ + { + kind: 'irrelevant', + reason: 'aggregator_over_official', + knownSources: ['https://example.com/official'], + source: 'web', + position: 1, + basis: 'output', + detail: + 'The official reference should rank before this aggregator.', + }, + ], + }, + }, + }); + assert.equal(response.status, 200); + const result = JSON.parse( + parseSseJson(await response.text()).result.content[0].text + ); + assert.equal(result.success, status === 200); + if (status !== 200) { + assert.equal(result.feedbackErrorCode, body.feedbackErrorCode); + assert.equal(result.retryable, status >= 500 || status === 429); + if (status === 429) assert.equal(result.retry_after_seconds, 60); + } + const submission = backend.requests.find( + (req) => req.url === '/v2/feedback' + ); + assert.ok(submission); + assert.equal(submission.body.endpoint, 'search'); + assert.equal( + submission.body.jobId, + '00000000-0000-4000-8000-000000000000' + ); + assert.equal(submission.body.rating, 'partial'); + assert.equal(submission.body.task, 'Read the API retry reference'); + assert.equal( + submission.body.assessment, + 'The reference explains supported retry intervals.' + ); + assert.equal(submission.headers.authorization, undefined); + assert.equal( + submission.headers['x-firecrawl-keyless-ip'], + '203.0.113.71' + ); + assert.equal( + submission.headers['x-firecrawl-keyless-secret'], + 'feedback-test-secret' + ); + assert.equal(submission.body.observations[0].position, 1); + assert.deepEqual(submission.body.observations[0].knownSources, [ + 'https://example.com/official', + ]); + if (status === 200) { + for (const endpoint of ['scrape', 'parse']) { + const args = { + endpoint, + jobId: '00000000-0000-4000-8000-000000000000', + rating: 'partial', + task: 'Extract the documented retry interval', + assessment: 'The structured output omitted the retry interval.', + ...(endpoint === 'parse' ? {docClass: 'born_digital'} : {}), + observations: [{kind: 'incorrect', reason: 'missing_fields', format: 'json', + basis: 'output', detail: 'The returned object has no retry interval.', + ...(endpoint === 'parse' ? {page: 2} : {location: 'Retry section'})}], + }; + const call = await httpToolCall(port, {id: `feedback-${endpoint}`, headers: {'x-forwarded-for': '203.0.113.71'}, params: {name: 'firecrawl_feedback', arguments: args}}); + assert.equal(JSON.parse(parseSseJson(await call.text()).result.content[0].text).success, true); + const posted = backend.requests.filter(req => req.url === '/v2/feedback').at(-1).body; + for (const field of ['endpoint', 'jobId', 'rating', 'task', 'assessment']) { + assert.equal(posted[field], args[field]); + } + assert.deepEqual(posted.observations, args.observations); + assert.equal(posted.docClass, args.docClass); + } + } + + if (status === 200) { + for (const field of ['task', 'assessment', 'observations', 'docClass']) { + const args = { endpoint: 'parse', jobId: '00000000-0000-4000-8000-000000000000', rating: 'partial', task: 'Read the retry reference', assessment: 'The retry interval was omitted from the output.', observations: [{kind: 'correct', basis: 'output', detail: 'The output contains the retry heading.'}], docClass: 'unknown' }; + delete args[field]; + const before = backend.requests.length; + const response = await httpToolCall(port, { id: `missing-${field}`, headers: {'x-forwarded-for': '203.0.113.71'}, params: {name: 'firecrawl_feedback', arguments: args} }); + const result = parseSseJson(await response.text()).result; + assert.equal(result.isError, true); + assert.match(result.content[0].text, new RegExp(`requires.*${field}`)); + assert.equal(backend.requests.length, before); + } + } + + assert.equal( + backend.requests.some((req) => req.url === '/v2/keyless/eligibility'), + false + ); + }); + } +}); + +for (const disabled of [false, true]) { + test(`keyless Search preserves references and retains invitations regardless of client feedback flags: ${disabled}`, async (t) => { + const metadata = { + jobId: '00000000-0000-4000-8000-000000000000', + feedback: { + endpoint: 'search', + message: 'Optional feedback is available.', + }, + }; + const backend = await startFakeFirecrawlBackend({ + keylessEligible: true, + searchResponse: { + status: 200, + body: { success: true, id: metadata.jobId, data: { web: [] }, metadata }, + }, + }); + t.after(() => backend.close()); + const port = await getFreePort(); + const child = spawnServer({ + CLOUD_SERVICE: 'true', + FIRECRAWL_NO_ENDPOINT_FEEDBACK: disabled ? 'true' : '', + FASTMCP_ENDPOINT: '/v2/mcp', + FIRECRAWL_API_URL: backend.url, + FIRECRAWL_OAUTH_ISSUER: backend.url, + HTTP_STREAMABLE_SERVER: 'true', + PORT: String(port), + KEYLESS_PROXY_SECRET: 'feedback-test-secret', + }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + for (const authenticated of [false, true]) { + const listed = await fetch(`http://127.0.0.1:${port}/v2/mcp`, { + method: 'POST', + headers: { + accept: 'application/json, text/event-stream', + 'content-type': 'application/json', + 'x-forwarded-for': '203.0.113.71', + ...(authenticated ? { Authorization: 'Bearer fc-feedback-test' } : {}), + }, + body: JSON.stringify({ id: 'feedback-tools', jsonrpc: '2.0', method: 'tools/list', params: {} }), + }); + const names = parseSseJson(await listed.text()).result.tools.map(tool => tool.name); + assert.equal(names.includes('firecrawl_feedback'), !authenticated || !disabled); + assert.equal(names.includes('firecrawl_search_feedback'), authenticated); + } + if (disabled) { + const blocked = await httpToolCall(port, { + id: 'disabled-authenticated-feedback', + headers: { Authorization: 'Bearer fc-feedback-test' }, + params: { name: 'firecrawl_feedback', arguments: { + endpoint: 'search', jobId: metadata.jobId, rating: 'good', + } }, + }); + assert.equal(parseSseJson(await blocked.text()).result.isError, true); + assert.equal(backend.requests.some(req => req.url === '/v2/feedback'), false); + } + const response = await httpToolCall(port, { + id: 'feedback-invitation', + headers: { 'x-forwarded-for': '203.0.113.71' }, + params: { + name: 'firecrawl_search', + arguments: { query: 'retry behavior' }, + }, + }); + const result = parseSseJson(await response.text()).result; + assert.deepEqual(JSON.parse(result.content[0].text).metadata, metadata); + assert.deepEqual(result.structuredContent.metadata, metadata); + assert.equal(result.structuredContent.id, metadata.jobId); + const call = backend.requests.find((req) => req.url === '/v2/search'); + assert.equal( + call.headers['x-firecrawl-no-feedback'], + undefined + ); + }); +} + +test('local Parse preserves job evidence on success and failure and retains invitations regardless of client feedback flags', async (t) => { + const directory = await mkdtemp(join(tmpdir(), 'feedback-parse-')); + t.after(() => rm(directory, { recursive: true, force: true })); + const filePath = join(directory, 'fixture.html'); + await writeFile(filePath, '

Parsed fixture

'); + for (const [status, success] of [[200, true], [500, false], [200, false]]) { + for (const disabled of [false, true]) { + await t.test(`HTTP ${status}, success ${success}, disabled ${disabled}`, async (t) => { + const metadata = { + jobId: '00000000-0000-4000-8000-000000000000', + feedback: { + endpoint: 'parse', + message: 'Optional feedback is available.', + }, + }; + const body = + success + ? { + success: true, + data: { markdown: '# Parsed fixture', metadata }, + } + : { success: false, error: 'Parsing failed', metadata }; + const backend = await startFakeFirecrawlBackend({ + parseResponse: { status, body }, + }); + t.after(() => backend.close()); + const child = spawnServer({ + FIRECRAWL_API_KEY: '', + FIRECRAWL_API_URL: backend.url, + CLOUD_SERVICE: '', + FIRECRAWL_DISABLE_ENDPOINT_FEEDBACK: disabled ? 'true' : '', + }); + t.after(() => stopChild(child)); + const client = new StdioMcpClient(child); + await client.request('initialize', { + capabilities: {}, + clientInfo: { name: 'parse-feedback-test', version: '0.0.0' }, + protocolVersion: '2025-06-18', + }); + client.notify('notifications/initialized'); + const listed = await client.request('tools/list', {}); + assert.equal(listed.tools.some(tool => tool.name === 'firecrawl_feedback'), true); + const result = await client.request('tools/call', { + name: 'firecrawl_parse', + arguments: { filePath }, + }); + const payload = + result.structuredContent ?? JSON.parse(result.content[0].text); + const returnedMetadata = payload.data?.metadata ?? payload.metadata; + assert.equal(returnedMetadata.jobId, metadata.jobId); + assert.equal(Boolean(returnedMetadata.feedback), true); + assert.equal(result.isError === true, !success); + if (!success) { + const textPayload = JSON.parse(result.content[0].text); + assert.equal(textPayload.error, 'Parsing failed'); + assert.equal(textPayload.metadata.jobId, metadata.jobId); + assert.equal(payload.error, 'Parsing failed'); + } + const call = backend.requests.find((req) => req.url === '/v2/parse'); + assert.equal(call.headers.authorization, undefined); + assert.equal( + call.headers['x-firecrawl-no-feedback'], + undefined + ); + assert.match(call.headers['content-type'], /multipart\/form-data/); + assert.match(call.raw, /Parsed fixture/); + }); + } + } +}); + + +test('local Parse requires an explicit API URL before reading or uploading files', async (t) => { + const directory = await mkdtemp(join(tmpdir(), 'parse-api-configuration-')); + t.after(() => rm(directory, { recursive: true, force: true })); + const filePath = join(directory, 'fixture.html'); + await writeFile(filePath, '

Controlled parse fixture

'); + for (const authenticated of [false, true]) { + await t.test(`authenticated ${authenticated}`, async (t) => { + const backend = await startFakeFirecrawlBackend(); + t.after(() => backend.close()); + const preloadPath = join(directory, `guard-${authenticated}.mjs`); + await writeFile(preloadPath, ` + import fs from 'node:fs/promises'; + import { syncBuiltinESMExports } from 'node:module'; + import { resolve } from 'node:path'; + const originalReadFile = fs.readFile; + fs.readFile = (...args) => { + if (typeof args[0] === 'string' && resolve(args[0]) === ${JSON.stringify(filePath)}) { + throw new Error('Test detected an unconfigured local file read'); + } + return originalReadFile(...args); + }; + syncBuiltinESMExports(); + const originalFetch = globalThis.fetch; + globalThis.fetch = (url, init) => { + if (url !== 'https://api.firecrawl.dev/v2/parse') { + throw new Error('Unexpected outbound URL'); + } + return originalFetch(${JSON.stringify(`${backend.url}/v2/parse`)}, init); + }; + `); + const child = spawnServer({ + FIRECRAWL_API_KEY: authenticated ? 'fc-parse-test' : '', + FIRECRAWL_OAUTH_TOKEN: '', + FIRECRAWL_API_URL: '', + CLOUD_SERVICE: '', + NODE_OPTIONS: `--import=${preloadPath}`, + }); + t.after(() => stopChild(child)); + const client = new StdioMcpClient(child); + await client.request('initialize', { + capabilities: {}, + clientInfo: { name: 'parse-api-configuration-test', version: '0.0.0' }, + protocolVersion: '2025-06-18', + }); + client.notify('notifications/initialized'); + for (const candidate of [filePath, relative(process.cwd(), filePath), join(directory, 'missing.html')]) { + const result = await client.request('tools/call', { + name: 'firecrawl_parse', + arguments: { filePath: candidate, formats: ['markdown'] }, + }); + assert.equal(result.isError, true); + assert.match(result.content[0].text, /requires.*FIRECRAWL_API_URL/); + assert.doesNotMatch(result.content[0].text, /unconfigured local file read|ENOENT/); + } + assert.equal(backend.requests.length, 0); + }); + } +}); + +for (const endpoint of ['search', 'scrape', 'parse']) { + for (const disabled of [false, true]) { + test(`authenticated ${endpoint} preserves references and leaves response metadata unchanged by feedback flags: ${disabled}`, async (t) => { + const metadata = { jobId: '00000000-0000-4000-8000-000000000000', feedback: { endpoint, message: 'Optional feedback.' } }; + const body = endpoint === 'search' + ? { success: true, id: metadata.jobId, data: { web: [] }, metadata } + : { success: true, data: { markdown: 'Observed content', metadata } }; + const backend = await startFakeFirecrawlBackend({ [`${endpoint}Response`]: { status: 200, body } }); + t.after(() => backend.close()); + const port = await getFreePort(); + const child = spawnServer({ CLOUD_SERVICE: 'true', FIRECRAWL_NO_ENDPOINT_FEEDBACK: disabled ? 'true' : '', + FASTMCP_ENDPOINT: '/v2/mcp', FIRECRAWL_API_URL: backend.url, FIRECRAWL_OAUTH_ISSUER: backend.url, + HTTP_STREAMABLE_SERVER: 'true', PORT: String(port), KEYLESS_PROXY_SECRET: 'feedback-test-secret' }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + const response = await httpToolCall(port, { id: 'authenticated-feedback-preference', headers: { Authorization: 'Bearer fc-feedback-test' }, params: { + name: `firecrawl_${endpoint}`, arguments: endpoint === 'search' ? { query: 'retry behavior' } : endpoint === 'scrape' ? { url: 'https://example.com/' } : { uploadRef: 'test-upload-ref' }, + } }); + const result = parseSseJson(await response.text()).result; + assert.notEqual(result.isError, true); + const payload = JSON.parse(result.content[0].text); + assert.deepEqual(payload.metadata ?? payload.data?.metadata, metadata); + const call = backend.requests.find(req => req.url === `/v2/${endpoint}`); + assert.equal(call.headers.authorization, 'Bearer fc-feedback-test'); + assert.equal(call.headers['x-firecrawl-no-feedback'], undefined); + }); + } +} + + +test('keyless Search failure preserves the feedback reference', async (t) => { + const metadata = { + jobId: '00000000-0000-4000-8000-000000000000', + feedback: { endpoint: 'search', docs: 'https://docs.firecrawl.dev/api-reference/endpoint/feedback' }, + }; + const backend = await startFakeFirecrawlBackend({ + keylessEligible: true, + searchResponse: { status: 200, body: { success: false, error: 'Search transport failed', metadata } }, + }); + t.after(() => backend.close()); + const port = await getFreePort(); + const child = spawnServer({ + CLOUD_SERVICE: 'true', FASTMCP_ENDPOINT: '/v2/mcp', + FIRECRAWL_API_URL: backend.url, FIRECRAWL_OAUTH_ISSUER: backend.url, + HTTP_STREAMABLE_SERVER: 'true', PORT: String(port), KEYLESS_PROXY_SECRET: 'feedback-test-secret', + }); + t.after(() => stopChild(child)); + await waitForHealth(port, child); + const response = await httpToolCall(port, { + id: 'failed-search-feedback', headers: { 'x-forwarded-for': '203.0.113.71' }, + params: { name: 'firecrawl_search', arguments: { query: 'retry reference' } }, + }); + const result = parseSseJson(await response.text()).result; + assert.equal(result.isError, true); + assert.deepEqual(result.structuredContent.metadata, metadata); + assert.equal(JSON.parse(result.content[0].text).metadata.jobId, metadata.jobId); +}); + test('every listed tool declares an output schema and returns structured content', async (t) => { const fakeApi = await startFakeFirecrawlApi(); t.after(() => fakeApi.close());