Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -419,7 +419,7 @@ Scrape content from a single URL with advanced options.

**Branding format:** Extracts comprehensive brand identity (colors, fonts, typography, spacing, logo, UI components) for design analysis or style replication.
**Privacy:** Set `redactPII: true` to return content with personally identifiable information redacted.
**Hosted server:** On the hosted server (`CLOUD_SERVICE=true`) scrape is read-only. It takes no browser `actions` and cannot accept provider terms. A named `profile` loads saved browser state without saving changes to it. An organization admin accepts terms in the dashboard.
**Hosted server:** On the hosted server (`CLOUD_SERVICE=true`) scrape is read-only. It takes no browser `actions` and cannot accept provider terms. A named `profile` loads saved browser state without saving changes to it, and the full endpoint's `firecrawl_search` treats `scrapeOptions.profile` the same way. To save browser state to a profile, open the page with `firecrawl_interact` (see below). An organization admin accepts terms in the dashboard.

**Returns:**

Expand Down Expand Up @@ -867,6 +867,7 @@ Interact with a fresh URL or with a page that was already opened by `firecrawl_s

- Pass `url` to scrape and open a page for interaction in one MCP call.
- Pass `scrapeId` to continue interacting with an existing scraped page.
- To save browser state (cookies, localStorage) to a named profile, pass `url` with `scrapeOptions: { "profile": { "name": "my-profile", "saveChanges": true } }`. The state is saved when `firecrawl_interact_stop` ends the session.
- Pass exactly one of `url` or `scrapeId`, plus either `prompt` or `code`.

**Usage Example:**
Expand Down
3 changes: 2 additions & 1 deletion docs/search-profile.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,7 +70,8 @@ same way. Alexandria results on this surface carry no `feedbackTool` pointer,
since `firecrawl_feedback` is not registered here.

`firecrawl_scrape` is read-only here (`readOnlyHint: true`): the surface runs in
hosted safe mode, so it takes no browser `actions`. Provider terms can be read
hosted safe mode, so it takes no browser `actions`, and a named `profile` loads
saved browser state without saving changes to it. Provider terms can be read
with the nested `terms/show` capability. As on the full surface, an organization
admin accepts them in the dashboard: `firecrawl_scrape` refuses every other
`terms/*` capability, and terms errors link to `requiresAction.url` or
Expand Down
33 changes: 26 additions & 7 deletions src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2268,13 +2268,31 @@ const scrapeParamsSchema = z.object({
.optional(),
});

// In safe mode firecrawl_scrape and firecrawl_search are read-only, so a named
// profile they open loads saved browser state without writing it back. The API
// saves profile changes unless told otherwise, so saveChanges: false is sent
// explicitly. firecrawl_interact (url with scrapeOptions.profile) saves them.
const readOnlyProfileSchema = z
.object({ name: z.string() })
.describe('Loads a saved browser profile without saving changes to it.');

function withReadOnlyProfile(
options: Record<string, unknown>
): Record<string, unknown> {
const profile = options.profile as { name: string } | undefined;
return SAFE_MODE && profile
? { ...options, profile: { name: profile.name, saveChanges: false } }
: options;
}

// firecrawl_scrape accepts either a page URL or an Exchange batch. The base
// schema stays url-required because search, crawl, and monitor reuse it for
// nested scrapeOptions, where `alexandria` has no meaning.
const ALEXANDRIA_IGNORED_SCRAPE_OPTIONS = new Set(['toolDetail', 'domainTools']);

const scrapeToolParamsSchema = scrapeParamsSchema
.extend({
...(SAFE_MODE ? { profile: readOnlyProfileSchema.optional() } : {}),
url: z.string().url().optional(),
timeout: z.number().int().positive().optional().describe("Execution timeout in milliseconds."),
requestId: z
Expand Down Expand Up @@ -2720,16 +2738,16 @@ const scrapeTool: RegisteredTool = {
name: 'firecrawl_scrape',
annotations: {
title: 'Firecrawl scrape',
// Hosted scrape omits browser actions and refuses provider terms writes
// before execution.
// Hosted scrape omits browser actions, loads profiles without saving,
// and refuses provider terms writes before execution.
readOnlyHint: SAFE_MODE,
openWorldHint: true, // Accepts any user-supplied URL on the public web.
destructiveHint: false, // Does not modify, delete, or write to external websites.
},
description: `
Scrape one URL and return its content: markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema. Use it when the request identifies a page and needs its content or defined fields. Use \`firecrawl_search\` when additional web sources are needed; on an authenticated session, \`firecrawl_map\` lists a site's URLs and \`firecrawl_crawl\` collects a set of pages.

Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch or a smaller \`maxAge\` to bound staleness. A successful response does not by itself confirm the page is still current. Browser actions can change the live page when interactive actions are enabled. Authenticated responses can include a \`metadata.scrapeId\` for optional scrape feedback.
Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch or a smaller \`maxAge\` to bound staleness. A successful response does not by itself confirm the page is still current. ${SAFE_MODE ? 'A named browser profile loads saved session data without saving changes to it.' : 'Browser actions can change the live page when interactive actions are enabled.'} Authenticated responses can include a \`metadata.scrapeId\` for optional scrape feedback.

On an authenticated session with Alexandria access, \`firecrawl_search\` with \`sources\` unset and \`firecrawl_find_tools\` can discover providers for the same fields across several pages; a matching provider returns typed records in one call. Keyless sessions have no provider matches.

Expand Down Expand Up @@ -2764,7 +2782,7 @@ Alexandria mode, on an authenticated session with Alexandria access: \`alexandri
const transformed = transformScrapeParams(
options as Record<string, unknown>
);
const cleaned = removeEmptyTopLevel(transformed);
const cleaned = withReadOnlyProfile(removeEmptyTopLevel(transformed));
if (cleaned.lockdown) {
log.info('Scraping URL (lockdown)');
} else {
Expand Down Expand Up @@ -2868,6 +2886,7 @@ For a programming question, add \`categories: ["developer"]\`; its hits return i
...searchToolBaseFields,
scrapeOptions: scrapeParamsSchema
.omit({ url: true })
.extend(SAFE_MODE ? { profile: readOnlyProfileSchema } : {})
.partial()
.optional()
.describe('Attach page content for web results in the same call. These fetches ignore maxAge, so use firecrawl_scrape when you need a live fetch. scrapeOptions fetches web pages, never Alexandria provider tools.'),
Expand All @@ -2890,8 +2909,8 @@ For a programming question, add \`categories: ["developer"]\`; its hits return i
searchOpts.toolDetail ??= 'compact';

if (searchOpts.scrapeOptions) {
searchOpts.scrapeOptions = transformScrapeParams(
searchOpts.scrapeOptions as Record<string, unknown>
searchOpts.scrapeOptions = withReadOnlyProfile(
transformScrapeParams(searchOpts.scrapeOptions as Record<string, unknown>)
);
}

Expand Down Expand Up @@ -3035,7 +3054,7 @@ server.addTool(findToolsTool);
// (no crawl, map, interact, monitor, parse or feedback references). Registered
// on the search surface in place of the module-level tools above.
const SEARCH_SURFACE_SCRAPE_DESCRIPTION = `
Scrape one URL and return its content, or execute catalogued Alexandria capabilities. URL mode returns markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema, plus page metadata. Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch. A successful response does not by itself confirm the page is still current. A named browser profile loads saved session data.
Scrape one URL and return its content, or execute catalogued Alexandria capabilities. URL mode returns markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema, plus page metadata. Firecrawl may serve recently indexed content; set \`maxAge: 0\` for a live fetch. A successful response does not by itself confirm the page is still current. ${SAFE_MODE ? 'A named browser profile loads saved session data without saving changes to it.' : 'Browser actions can change the live page, and a named browser profile can load saved session data and overwrite its stored state.'}

\`firecrawl_search\` with \`sources\` unset and \`firecrawl_find_tools\` can discover providers for the same fields across several pages; a matching Alexandria provider returns typed records in one call.

Expand Down
16 changes: 12 additions & 4 deletions tests/mcp-read-only-scrape.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ async function startHosted(t) {
return { api, port, searchPort };
}

test('hosted scrape is read-only and preserves existing profile and search options', async (t) => {
test('hosted scrape and search are read-only and load profiles without saving', async (t) => {
const { api, port, searchPort } = await startHosted(t);
const headers = { 'x-api-key': 'fc-hosted-test' };
for (const [label, endpoint, surfacePort] of [
Expand All @@ -78,13 +78,20 @@ test('hosted scrape is read-only and preserves existing profile and search optio
assert.equal(scrape.annotations.readOnlyHint, true, label);
assert.equal(scrape.annotations.destructiveHint, false, label);
assert.equal(scrape.inputSchema.properties.actions, undefined, `${label}: no browser actions`);
assert.doesNotMatch(scrape.description, /overwrite its stored state/, label);
assert.doesNotMatch(scrape.description, /overwrite its stored state|Browser actions/, label);
assert.match(scrape.description, /without saving changes/, label);
assert.deepEqual(Object.keys(scrape.inputSchema.properties.profile.properties), ['name'], `${label}: profile takes only a name`);
}

const { tools } = await httpSession(port, '/v2/mcp', headers);
const byName = new Map(tools.map((tool) => [tool.name, tool]));
assert.equal(byName.get('firecrawl_agent').annotations.readOnlyHint, false);
assert.equal(byName.get('firecrawl_search').annotations.readOnlyHint, true);
assert.deepEqual(
Object.keys(byName.get('firecrawl_search').inputSchema.properties.scrapeOptions.properties.profile.properties),
['name'],
'search scrapeOptions profile takes only a name'
);
assert.equal(tools.some((tool) => /terms/.test(tool.name)), false, 'no tool accepts terms');

for (const [surfacePort, endpoint] of [[port, '/v2/mcp'], [searchPort, '/v2/mcp-search']]) {
Expand All @@ -99,7 +106,8 @@ test('hosted scrape is read-only and preserves existing profile and search optio
headers,
});
assert.notEqual(scraped.isError, true, JSON.stringify(scraped));
assert.deepEqual(api.requests.at(-1).body.profile, profile, 'profile options remain unchanged');
// A call shaped by an older tool definition may still send saveChanges.
assert.deepEqual(api.requests.at(-1).body.profile, { name: 'saved-login', saveChanges: false }, 'profile never saves');
assert.equal(api.requests.at(-1).body.actions, undefined);
}
}
Expand All @@ -114,7 +122,7 @@ test('hosted scrape is read-only and preserves existing profile and search optio
headers,
});
assert.notEqual(searched.isError, true, JSON.stringify(searched));
assert.deepEqual(api.requests.at(-1).body.scrapeOptions.profile, { name: 'saved-login', saveChanges: true });
assert.deepEqual(api.requests.at(-1).body.scrapeOptions.profile, { name: 'saved-login', saveChanges: false });
assert.equal(api.requests.at(-1).body.scrapeOptions.actions, undefined);

const { client } = await startStdio(t, { CLOUD_SERVICE: 'false', FIRECRAWL_API_KEY: 'fc-test', FIRECRAWL_API_URL: api.url });
Expand Down
Loading