From 53020decd438cb1f442c1d2af600e62e59f85f88 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 24 Sep 2026 13:18:35 +0000 Subject: [PATCH] docs: add agent quickstart files for all SDK languages Add one canonical quickstart per SDK language in agent-quickstart/: - node.mdx (JS/TS SDK v4.41.0) - python.mdx (Python SDK v4.44.0) - rust.mdx (Rust SDK v2.21.0) - java.mdx (Java SDK v1.18.0) - elixir.mdx (Elixir SDK v1.11.0) Each file covers search, scrape, and interact with confirmed parameters, working examples, and language-specific notes. Generated from SDK source as primary authority, cross-referenced with the v2 OpenAPI spec. Co-Authored-By: Claude Opus 4.6 (1M context) Claude-Session: https://claude.ai/code/session_01XyF46EHnn2GAewQjxb12LU --- agent-quickstart/elixir.mdx | 199 +++++++++++++++++++++++++++++ agent-quickstart/java.mdx | 229 +++++++++++++++++++++++++++++++++ agent-quickstart/node.mdx | 202 +++++++++++++++++++++++++++++ agent-quickstart/python.mdx | 198 +++++++++++++++++++++++++++++ agent-quickstart/rust.mdx | 245 ++++++++++++++++++++++++++++++++++++ 5 files changed, 1073 insertions(+) create mode 100644 agent-quickstart/elixir.mdx create mode 100644 agent-quickstart/java.mdx create mode 100644 agent-quickstart/node.mdx create mode 100644 agent-quickstart/python.mdx create mode 100644 agent-quickstart/rust.mdx diff --git a/agent-quickstart/elixir.mdx b/agent-quickstart/elixir.mdx new file mode 100644 index 000000000..2cae876d6 --- /dev/null +++ b/agent-quickstart/elixir.mdx @@ -0,0 +1,199 @@ +--- +title: "Elixir Agent Quickstart" +description: "Canonical Firecrawl Elixir quickstart for external agents using search, scrape, and interact." +--- + +# Firecrawl Elixir Agent Quickstart + +Canonical quickstart for external agents. Generated from SDK source (`:firecrawl` **v1.11.0**) and the v2 OpenAPI spec. + +## Install + +Add to `mix.exs`: + +```elixir +{:firecrawl, "~> 1.11.0"} +``` + +## Authenticate + +```elixir +# config/runtime.exs or config.exs +config :firecrawl, api_key: System.get_env("FIRECRAWL_API_KEY") + +# Or pass api_key per call: +{:ok, res} = Firecrawl.search_and_scrape( + [query: "site:docs.firecrawl.dev webhook retries"], + api_key: "fc-your-api-key" +) +``` + +Every function accepts a trailing `opts` keyword list supporting `:api_key` (override the global key) and `:base_url` (override the default `https://api.firecrawl.dev/v2`). Omitting `api_key` uses the keyless free tier. + +## When To Use What + +- **search** — start with a query and need discovery. Returns ranked results you can then scrape. +- **scrape** — you already have a URL and want page content in one or more formats. +- **interact** — the page needs clicks, forms, or post-scrape browser actions. Requires a scrape job ID from a prior scrape. + +## Search + +### Why use it + +Discover relevant pages from a query. Constrain to a site with `site:` in the query string (e.g. `site:docs.firecrawl.dev crawl webhooks`). + +### Preferred SDK method + +`Firecrawl.search_and_scrape(params \\ [], opts \\ [])` → `{:ok, Req.Response.t()} | {:error, Exception.t()}` + +### Example + +```elixir +{:ok, res} = Firecrawl.search_and_scrape( + query: "site:docs.firecrawl.dev webhook retries", + limit: 5, + scrape_options: [ + formats: ["markdown"], + only_main_content: true + ] +) + +web_results = res.body["data"]["web"] +``` + +### Parameters + +All parameters are keyword list entries. The SDK validates them at compile time via NimbleOptions and converts snake_case keys to camelCase JSON. + +| Parameter | Type | Description | +|---|---|---| +| `query` | `string` | The search query (required). Use `site:example.com` to scope to a domain. | +| `sources` | `list(atom \| string \| map)` | Which sources: `:web`, `:news`, `:images` (or string/map equivalents). | +| `categories` | `list(any)` | Filter by category: `:developer`, `:research`, `:pdf`. | +| `include_domains` | `list(string)` | Only include results from these domains. | +| `exclude_domains` | `list(string)` | Exclude results from these domains. | +| `limit` | `integer` | Cap on number of results. | +| `tbs` | `string` | Time-based filter (e.g. `"qdr:d"`, `"qdr:w"`). | +| `location` | `string` | Location string for localized results. | +| `country` | `string` | ISO country code for geo-targeting. | +| `ignore_invalid_urls` | `boolean` | Drop URLs that cannot be scraped. | +| `timeout` | `integer` | Request timeout in milliseconds. | +| `highlights` | `boolean` | Generate query-relevant highlights. | +| `scrape_options` | `keyword` | Scrape each search result (see Scrape parameters). | +| `enterprise` | `list(string)` | Enterprise options: `"zdr"`, `"anon"`. | + +## Scrape + +### Why use it + +Get structured content from a URL in one or more formats — markdown, HTML, JSON extraction, screenshots, and more. + +### Preferred SDK method + +`Firecrawl.scrape_and_extract_from_url(params \\ [], opts \\ [])` → `{:ok, Req.Response.t()} | {:error, Exception.t()}` + +### Example + +```elixir +{:ok, res} = Firecrawl.scrape_and_extract_from_url( + url: "https://example.com/pricing", + formats: [ + "markdown", + %{type: "json", prompt: "Extract plan names and prices."} + ], + only_main_content: true +) + +markdown = res.body["data"]["markdown"] +json_data = res.body["data"]["json"] +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `url` | `string` | The URL to scrape (required). | +| `formats` | `list(string \| map)` | Output formats. Strings: `"markdown"`, `"html"`, `"rawHtml"`, `"links"`, `"images"`, `"screenshot"`, `"summary"`, `"changeTracking"`, `"json"`, `"branding"`, `"audio"`, `"video"`. Maps: `%{type: "json", prompt: ..., schema: ...}`, `%{type: "question", question: ...}`, `%{type: "highlights", query: ...}`, `%{type: "screenshot", fullPage: ..., quality: ..., viewport: ...}`, `%{type: "changeTracking", modes: [...]}`, `%{type: "attributes", selectors: [...]}`. | +| `headers` | `map` | Custom request headers. | +| `include_tags` | `list(string)` | Only include content from these HTML tags. | +| `exclude_tags` | `list(string)` | Exclude content from these HTML tags. | +| `only_main_content` | `boolean` | Strip nav, footer, and boilerplate. | +| `timeout` | `integer` | Timeout in milliseconds. | +| `wait_for` | `integer` | Wait for the page to render (milliseconds). | +| `mobile` | `boolean` | Emulate a mobile device. | +| `parsers` | `list(string \| map)` | File parsing controls (e.g. `"pdf"` or `%{type: "pdf", mode: "auto", maxPages: 5}`). | +| `actions` | `list(map)` | Pre-scrape browser actions. Types: `wait`, `click`, `write`, `press`, `scroll`, `screenshot`, `scrape`, `executeJavascript`, `pdf`. | +| `location` | `keyword` | Geo-aware scraping (e.g. `[country: "US", languages: ["en-US"]]`). | +| `skip_tls_verification` | `boolean` | Skip TLS verification. | +| `remove_base64_images` | `boolean` | Drop base64 images from markdown. | +| `block_ads` | `boolean` | Block ads and cookie popups. | +| `proxy` | `atom \| string` | Proxy control: `:basic`, `:enhanced`, `:auto`. | +| `max_age` | `integer` | Use cached data up to this age (milliseconds). | +| `min_age` | `integer` | Use cached data only if at least this old (milliseconds). | +| `store_in_cache` | `boolean` | Cache the result in Firecrawl. | +| `lockdown` | `boolean` | Only serve cached results, never make outbound requests. | +| `redact_pii` | `boolean` | Redact personally identifiable information. | +| `profile` | `keyword` | Persistent browser profile (`[name: "...", save_changes: true]`). | +| `zero_data_retention` | `boolean` | Enable zero data retention. | +| `audit_metadata` | `keyword` | User attribution for SIEM logging (`[username: "..."]`). | + +## Interact + +### Why use it + +Run code in the browser session tied to a scrape job. The Elixir SDK exposes code-based interactions only (no `prompt` parameter — use the Node.js, Python, or Rust SDK for natural-language browser control). + +### Preferred SDK method + +`Firecrawl.interact_with_scrape_browser_session(job_id, params \\ [], opts \\ [])` → `{:ok, Req.Response.t()} | {:error, Exception.t()}` + +### Example + +```elixir +{:ok, scrape_res} = Firecrawl.scrape_and_extract_from_url( + url: "https://example.com", + formats: ["markdown"] +) + +job_id = scrape_res.body["data"]["metadata"]["scrapeId"] + +{:ok, res} = Firecrawl.interact_with_scrape_browser_session( + job_id, + code: "console.log(await page.title());", + language: :node, + timeout: 60 +) + +IO.inspect(res.body) + +# Stop the session when done: +{:ok, _} = Firecrawl.stop_interactive_scrape_browser_session(job_id) +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `job_id` | `String.t()` | Scrape job ID from `data.metadata.scrapeId` (first positional argument). | +| `code` | `string` | Code to execute in the browser session (required). | +| `language` | `atom \| string` | Runtime: `:python`, `:node`, `:bash`. | +| `timeout` | `integer` | Execution timeout in seconds. | +| `origin` | `string` | Optional origin label for telemetry. | + +Stop session: `Firecrawl.stop_interactive_scrape_browser_session(job_id, opts \\ [])` — issues `DELETE /scrape/{jobId}/interact`. + +## Notes + +- The Elixir client is **auto-generated from the OpenAPI spec**. Function names and parameter keys are generated, not hand-written. +- All parameters use **snake_case keyword lists**. The SDK converts to camelCase JSON automatically. +- Every function has a **bang (`!`) variant** that raises on error instead of returning `{:error, _}` (e.g. `Firecrawl.scrape_and_extract_from_url!`). +- No struct-based client — a fresh `Req` HTTP client is built per request. +- SDK auto-injects `"origin": "elixir-sdk@1.11.0"` into all request bodies. +- Atom values (like `:node`, `:basic`) are automatically converted to strings before JSON serialization. +- Nested keyword lists are recursively converted to camelCased maps. + +## Source Of Truth + +- `firecrawl/apps/elixir-sdk/mix.exs` +- `firecrawl/apps/elixir-sdk/lib/firecrawl.ex` +- `firecrawl-docs/api-reference/v2-openapi.json` diff --git a/agent-quickstart/java.mdx b/agent-quickstart/java.mdx new file mode 100644 index 000000000..fbb85cf22 --- /dev/null +++ b/agent-quickstart/java.mdx @@ -0,0 +1,229 @@ +--- +title: "Java Agent Quickstart" +description: "Canonical Firecrawl Java quickstart for external agents using search, scrape, and interact." +--- + +# Firecrawl Java Agent Quickstart + +Canonical quickstart for external agents. Generated from SDK source (`firecrawl-java` **v1.18.0**) and the v2 OpenAPI spec. + +## Install + +Maven: + +```xml + + com.firecrawl + firecrawl-java + 1.18.0 + +``` + +Gradle: + +```gradle +implementation("com.firecrawl:firecrawl-java:1.18.0") +``` + +## Authenticate + +```java +import com.firecrawl.client.FirecrawlClient; + +FirecrawlClient client = FirecrawlClient.builder() + .apiKey(System.getenv("FIRECRAWL_API_KEY")) + .build(); + +// Or from environment (reads FIRECRAWL_API_KEY env / firecrawl.apiKey sysprop): +// FirecrawlClient client = FirecrawlClient.fromEnv(); +``` + +Builder options: `apiKey` (optional — omitting uses keyless free tier), `apiUrl` (default `https://api.firecrawl.dev`), `timeoutMs` (default `300000`), `maxRetries` (default `3`), `backoffFactor` (default `0.5`), `asyncExecutor`, `httpClient`. + +## When To Use What + +- **search** — start with a query and need discovery. Returns ranked results you can then scrape. +- **scrape** — you already have a URL and want page content in one or more formats. +- **interact** — the page needs clicks, forms, or post-scrape browser actions. Requires a scrape job ID from a prior scrape. + +## Search + +### Why use it + +Discover relevant pages from a query. Constrain to a site with `site:` in the query string (e.g. `site:docs.firecrawl.dev crawl webhooks`). + +### Preferred SDK method + +`client.search(query)` or `client.search(query, options)` → `SearchData` + +### Example + +```java +import com.firecrawl.models.SearchOptions; +import com.firecrawl.models.ScrapeOptions; +import com.firecrawl.models.SearchData; +import java.util.List; +import java.util.Map; + +SearchData results = client.search( + "site:docs.firecrawl.dev webhook retries", + SearchOptions.builder() + .limit(5) + .scrapeOptions( + ScrapeOptions.builder() + .formats(List.of("markdown")) + .onlyMainContent(true) + .build() + ) + .build() +); + +List> web = results.getWeb(); +``` + +**Wrong turn to avoid:** `search()` returns `SearchData`. Access results via `getWeb()`, `getNews()`, `getImages()` — each returns `List>` (may be null). Do not treat `SearchData` as a directly iterable list. + +### Parameters + +All `SearchOptions` fields are nullable (omitted from JSON when null). Use `SearchOptions.builder()...build()`. + +| Parameter | Type | Description | +|---|---|---| +| `query` | `String` | The search query. Use `site:example.com` to scope to a domain. | +| `sources` | `List` | Which sources to search: `"web"`, `"news"`, `"images"`. | +| `categories` | `List` | Filter by category: `"github"`, `"research"`, `"pdf"`. | +| `includeDomains` | `List` | Only include results from these domains. | +| `excludeDomains` | `List` | Exclude results from these domains. | +| `limit` | `Integer` | Cap on number of results. | +| `tbs` | `String` | Time-based filter (e.g. `"qdr:d"`, `"qdr:w"`). | +| `location` | `String` | Location string for localized results. | +| `country` | `String` | ISO country code for geo-targeting. | +| `ignoreInvalidURLs` | `Boolean` | Drop URLs that cannot be scraped. | +| `timeout` | `Integer` | Request timeout in milliseconds. | +| `highlights` | `Boolean` | Generate query-relevant highlights. Defaults to `true`. | +| `scrapeOptions` | `ScrapeOptions` | Scrape each search result (see Scrape parameters). | +| `integration` | `String` | Integration identifier. | + +## Scrape + +### Why use it + +Get structured content from a URL in one or more formats — markdown, HTML, JSON extraction, screenshots, and more. + +### Preferred SDK method + +`client.scrape(url)` or `client.scrape(url, options)` → `Document` + +### Example + +```java +import com.firecrawl.models.ScrapeOptions; +import com.firecrawl.models.JsonFormat; +import com.firecrawl.models.Document; + +Document doc = client.scrape( + "https://example.com/pricing", + ScrapeOptions.builder() + .formats(List.of( + "markdown", + JsonFormat.builder() + .prompt("Extract plan names and prices.") + .build() + )) + .onlyMainContent(true) + .build() +); + +System.out.println(doc.getMarkdown()); +System.out.println(doc.getJson()); +``` + +### Parameters + +All `ScrapeOptions` fields are nullable. Use `ScrapeOptions.builder()...build()`. + +| Parameter | Type | Description | +|---|---|---| +| `url` | `String` | The URL to scrape. | +| `formats` | `List` | Output formats. Strings: `"markdown"`, `"html"`, `"rawHtml"`, `"links"`, `"images"`, `"screenshot"`, `"json"`, `"audio"`, `"video"`. Objects: `JsonFormat.builder().prompt(...).build()`, `QuestionFormat`, `HighlightsFormat`, maps for screenshot/changeTracking/attributes. | +| `headers` | `Map` | Custom request headers. | +| `includeTags` | `List` | Only include content from these HTML tags. | +| `excludeTags` | `List` | Exclude content from these HTML tags. | +| `onlyMainContent` | `Boolean` | Strip nav, footer, and boilerplate. | +| `timeout` | `Integer` | Timeout in milliseconds. | +| `waitFor` | `Integer` | Wait for the page to render (milliseconds). | +| `mobile` | `Boolean` | Emulate a mobile device. | +| `parsers` | `List` | File parsing controls (e.g. `"pdf"` or `PdfParser` with `maxPages`). | +| `actions` | `List>` | Pre-scrape browser actions. Types: `wait`, `click`, `write`, `press`, `scroll`, `screenshot`, `scrape`, `executeJavascript`, `pdf`. | +| `location` | `LocationConfig` | Geo-aware scraping. Built with `LocationConfig.builder().country("US").languages(List.of("en-US")).build()`. | +| `skipTlsVerification` | `Boolean` | Skip TLS verification. | +| `removeBase64Images` | `Boolean` | Drop base64 images from markdown. | +| `blockAds` | `Boolean` | Block ads and cookie popups. | +| `proxy` | `String` | Proxy control: `"basic"`, `"stealth"`, `"enhanced"`, `"auto"`, or custom URL. | +| `maxAge` | `Long` | Use cached data up to this age (milliseconds). | +| `storeInCache` | `Boolean` | Cache the result in Firecrawl. | +| `lockdown` | `Boolean` | Only serve cached results, never make outbound requests. | +| `redactPII` | `Boolean` | Redact personally identifiable information. | +| `auditMetadata` | `AuditMetadata` | User attribution for SIEM logging. | +| `integration` | `String` | Integration identifier. | + +## Interact + +### Why use it + +Run code in the browser session tied to a scrape job. The Java SDK exposes code-based interactions only (no `prompt` parameter — use the Node.js or Python SDK for natural-language browser control). + +### Preferred SDK method + +`client.interact(jobId, code)` or `client.interact(jobId, code, language, timeout)` → `BrowserExecuteResponse` + +### Example + +```java +import com.firecrawl.models.BrowserExecuteResponse; +import com.firecrawl.models.BrowserDeleteResponse; + +Document doc = client.scrape("https://example.com", + ScrapeOptions.builder().formats(List.of("markdown")).build()); +String jobId = (String) doc.getMetadata().get("scrapeId"); + +BrowserExecuteResponse result = client.interact( + jobId, + "console.log(await page.title());", + "node", + 60 +); +System.out.println(result.getStdout()); + +// Stop the session when done: +BrowserDeleteResponse stopped = client.stopInteractiveBrowser(jobId); +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `jobId` | `String` | Scrape job ID from `document.getMetadata().get("scrapeId")`. | +| `code` | `String` | Code to execute in the browser session. | +| `language` | `String` | Runtime: `"python"`, `"node"`, `"bash"`. Default: `"node"`. | +| `timeout` | `Integer` | Execution timeout in seconds (1–300). Null uses API default (30s). | +| `origin` | `String` | Optional origin label for request attribution (5-parameter overload only). | + +Stop session: `client.stopInteractiveBrowser(jobId)` → `BrowserDeleteResponse` + +## Notes + +- All parameter names use **camelCase**. +- Deprecated aliases: `scrapeExecute` → `interact`, `deleteScrapeBrowser` → `stopInteractiveBrowser`. Always use the preferred names. +- Async variants exist for all methods: `searchAsync`, `scrapeAsync`, `interactAsync`, `stopInteractiveBrowserAsync` — returning `CompletableFuture`. +- `ScrapeOptions.formats` and `SearchOptions.sources` are `List` — they accept both strings and typed objects (e.g. `JsonFormat`, `QuestionFormat`) in the same list. +- `SearchData` result lists are `List>` — individual fields are accessed as map keys. +- SDK auto-injects `"origin": "java-sdk@1.18.0"` into request bodies. + +## Source Of Truth + +- `firecrawl/apps/java-sdk/build.gradle.kts` +- `firecrawl/apps/java-sdk/src/main/java/com/firecrawl/client/FirecrawlClient.java` +- `firecrawl/apps/java-sdk/src/main/java/com/firecrawl/models/ScrapeOptions.java` +- `firecrawl/apps/java-sdk/src/main/java/com/firecrawl/models/SearchOptions.java` +- `firecrawl-docs/api-reference/v2-openapi.json` diff --git a/agent-quickstart/node.mdx b/agent-quickstart/node.mdx new file mode 100644 index 000000000..60927b9eb --- /dev/null +++ b/agent-quickstart/node.mdx @@ -0,0 +1,202 @@ +--- +title: "Node.js Agent Quickstart" +description: "Canonical Firecrawl Node.js quickstart for external agents using search, scrape, and interact." +--- + +# Firecrawl Node.js Agent Quickstart + +Canonical quickstart for external agents. Generated from SDK source (`firecrawl` **v4.41.0**) and the v2 OpenAPI spec. + +## Install + +```bash +npm install firecrawl +``` + +## Authenticate + +```ts +import Firecrawl from "firecrawl"; + +const client = new Firecrawl({ + apiKey: process.env.FIRECRAWL_API_KEY, +}); +``` + +Constructor options: `apiKey` (falls back to `FIRECRAWL_API_KEY` env), `apiUrl` (falls back to `FIRECRAWL_API_URL` env or `https://api.firecrawl.dev`), `timeoutMs`, `maxRetries`, `backoffFactor`. All optional — omitting `apiKey` uses the keyless free tier. + +## When To Use What + +- **search** — start with a query and need discovery. Returns ranked results you can then scrape. +- **scrape** — you already have a URL and want page content in one or more formats. +- **interact** — the page needs clicks, forms, or post-scrape browser actions. Requires a `scrapeId` from a prior scrape. + +## Search + +### Why use it + +Discover relevant pages from a query. Constrain to a site with `site:` in the query string (e.g. `site:docs.firecrawl.dev crawl webhooks`). + +### Preferred SDK method + +`client.search(query, options?)` → `Promise` + +### Example + +```ts +const results = await client.search("site:docs.firecrawl.dev webhook retries", { + limit: 5, + scrapeOptions: { + formats: ["markdown"], + onlyMainContent: true, + }, +}); + +for (const item of results.web ?? []) { + console.log(item.url, item.title); +} +``` + +**Wrong turn to avoid:** `search()` does not return `{ data: [...] }`. Access results via `result.web`, `result.news`, `result.images`, or `result.tools`. + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `query` | `string` | The search query. Use `site:example.com` to scope to a domain. | +| `sources` | `("web" \| "news" \| "images" \| "alexandria")[]` | Which sources to search. | +| `categories` | `("github" \| "research" \| "pdf" \| "developer")[]` | Filter results by category. | +| `includeDomains` | `string[]` | Only include results from these domains. Mutually exclusive with `excludeDomains`. | +| `excludeDomains` | `string[]` | Exclude results from these domains. Mutually exclusive with `includeDomains`. | +| `limit` | `number` | Cap on number of results. | +| `tbs` | `string` | Time-based filter (e.g. `qdr:d` for past day, `qdr:w` for past week). | +| `location` | `string` | Location string for localized results. | +| `country` | `string` | ISO 3166-1 alpha-2 country code for geo-targeting. | +| `ignoreInvalidURLs` | `boolean` | Drop URLs that cannot be scraped. | +| `timeout` | `number` | Request timeout in milliseconds. | +| `highlights` | `boolean` | Generate query-relevant highlights. Defaults to `true`. | +| `scrapeOptions` | `ScrapeOptions` | Scrape each search result with these options (see Scrape parameters). | +| `domainTools` | `boolean` | Include domain-matched Alexandria tools in results. | +| `toolDetail` | `"compact" \| "summary" \| "full"` | Detail level for discovered tools. | +| `enterprise` | `("default" \| "anon" \| "zdr")[]` | Enterprise options for zero data retention. | +| `threatProtection` | `ThreatProtectionOptions` | Enterprise threat protection settings. | +| `integration` | `string` | Integration identifier for server-side tracking. | + +## Scrape + +### Why use it + +Get structured content from a URL in one or more formats — markdown, HTML, JSON extraction, screenshots, and more. + +### Preferred SDK method + +`client.scrape(url, options?)` → `Promise` + +### Example + +```ts +const doc = await client.scrape("https://example.com/pricing", { + formats: [ + "markdown", + { type: "json", prompt: "Extract plan names and prices." }, + ], + onlyMainContent: true, +}); + +console.log(doc.markdown); +console.log(doc.json); +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `url` | `string` | The URL to scrape. | +| `formats` | `FormatOption[]` | Output formats. Strings: `"markdown"`, `"html"`, `"rawHtml"`, `"links"`, `"images"`, `"screenshot"`, `"summary"`, `"changeTracking"`, `"attributes"`, `"branding"`, `"product"`, `"menu"`, `"audio"`, `"video"`. Objects: `{ type: "json", prompt?, schema? }`, `{ type: "question", question }`, `{ type: "highlights", query }`, `{ type: "screenshot", fullPage?, quality?, viewport? }`, `{ type: "changeTracking", modes, schema?, prompt?, tag? }`, `{ type: "attributes", selectors: [{ selector, attribute }] }`. Note: plain `"json"` string is rejected — use the object form. | +| `headers` | `Record` | Custom request headers. | +| `includeTags` | `string[]` | Only include content from these HTML tags. | +| `excludeTags` | `string[]` | Exclude content from these HTML tags. | +| `onlyMainContent` | `boolean` | Strip nav, footer, and boilerplate. | +| `timeout` | `number` | Timeout in milliseconds. | +| `waitFor` | `number` | Wait for the page to render (milliseconds). | +| `mobile` | `boolean` | Emulate a mobile device. | +| `parsers` | `(string \| { type: "pdf", mode?, maxPages? })[]` | File parsing controls. | +| `actions` | `ActionOption[]` | Pre-scrape browser actions. Types: `wait` (milliseconds or selector), `click` (selector), `write` (text), `press` (key), `scroll` (direction: up/down), `screenshot`, `scrape`, `executeJavascript` (script), `pdf` (format, landscape, scale). | +| `location` | `{ country?: string, languages?: string[] }` | Geo and language-aware scraping. | +| `skipTlsVerification` | `boolean` | Skip TLS verification. | +| `removeBase64Images` | `boolean` | Drop base64 images from markdown. | +| `fastMode` | `boolean` | Faster scrapes with reduced fidelity. | +| `blockAds` | `boolean` | Block ads and cookie popups. | +| `proxy` | `"basic" \| "stealth" \| "enhanced" \| "auto" \| string` | Proxy control. | +| `maxAge` | `number` | Use cached data up to this age (milliseconds). | +| `minAge` | `number` | Use cached data only if at least this old (milliseconds). | +| `storeInCache` | `boolean` | Cache the result in Firecrawl. | +| `lockdown` | `boolean` | Only serve cached results, never make outbound requests. | +| `redactPII` | `boolean \| RedactPIIOptions` | Redact personally identifiable information. | +| `profile` | `{ name: string, saveChanges?: boolean }` | Persistent browser profile across scrapes and interactions. | +| `domainTools` | `boolean` | Include domain-matched Alexandria tools. | +| `toolDetail` | `"compact" \| "summary" \| "full"` | Detail level for discovered tools. | +| `auditMetadata` | `{ username: string }` | User attribution for SIEM logging. | +| `threatProtection` | `ThreatProtectionOptions` | Enterprise threat protection. | +| `integration` | `string` | Integration identifier. | + +## Interact + +### Why use it + +Control the browser session tied to a scrape job — run code or give natural-language instructions. Requires a `scrapeId` from a prior scrape's `metadata`. + +### Preferred SDK method + +`client.interact(jobId, args)` → `Promise` + +### Example + +```ts +const doc = await client.scrape("https://example.com", { + formats: ["markdown"], +}); +const jobId = doc.metadata?.scrapeId; + +const result = await client.interact(jobId, { + prompt: "Click the pricing tab and summarize the plans.", +}); +console.log(result.output); + +// Or run code directly: +const codeResult = await client.interact(jobId, { + code: "console.log(await page.title());", + language: "node", + timeout: 60, +}); +console.log(codeResult.stdout); + +// Stop the session when done: +await client.stopInteraction(jobId); +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `jobId` | `string` | Scrape job ID from `document.metadata.scrapeId`. | +| `code` | `string` | Code to execute in the browser session. At least one of `code` or `prompt` required. | +| `prompt` | `string` | Natural-language instruction for the browser agent. At least one of `code` or `prompt` required. | +| `language` | `"python" \| "node" \| "bash"` | Runtime for code execution. Default: `"node"`. | +| `timeout` | `number` | Execution timeout in seconds. | + +Stop session: `client.stopInteraction(jobId)` → `Promise` + +## Notes + +- All parameter names use **camelCase**. +- Deprecated aliases: `scrapeUrl` → `scrape`, `scrapeExecute` → `interact`, `stopInteractiveBrowser` / `deleteScrapeBrowser` → `stopInteraction`. Always use the preferred names. +- When `formats` includes `{ type: "json", schema: ZodSchema }`, the returned `json` field is typed to the Zod output type. +- `autoResume` (SDK-only, default `true`) controls SDK-side retry for long-running scrapes. Set `false` to surface timeout errors immediately. + +## Source Of Truth + +- `firecrawl/apps/js-sdk/firecrawl/src/index.ts` +- `firecrawl/apps/js-sdk/firecrawl/src/v2/client.ts` +- `firecrawl/apps/js-sdk/firecrawl/src/v2/types.ts` +- `firecrawl-docs/api-reference/v2-openapi.json` diff --git a/agent-quickstart/python.mdx b/agent-quickstart/python.mdx new file mode 100644 index 000000000..bae6224a2 --- /dev/null +++ b/agent-quickstart/python.mdx @@ -0,0 +1,198 @@ +--- +title: "Python Agent Quickstart" +description: "Canonical Firecrawl Python quickstart for external agents using search, scrape, and interact." +--- + +# Firecrawl Python Agent Quickstart + +Canonical quickstart for external agents. Generated from SDK source (`firecrawl-py` **v4.44.0**) and the v2 OpenAPI spec. + +## Install + +```bash +pip install firecrawl-py +``` + +## Authenticate + +```python +import os +from firecrawl import Firecrawl + +client = Firecrawl(api_key=os.environ.get("FIRECRAWL_API_KEY")) +``` + +Constructor parameters: `api_key` (falls back to `FIRECRAWL_API_KEY` env), `api_url` (default `https://api.firecrawl.dev`), `timeout` (seconds), `max_retries` (default `3`), `backoff_factor` (default `0.5`). All optional — omitting `api_key` uses the keyless free tier. An async variant is available as `AsyncFirecrawl`. + +## When To Use What + +- **search** — start with a query and need discovery. Returns ranked results you can then scrape. +- **scrape** — you already have a URL and want page content in one or more formats. +- **interact** — the page needs clicks, forms, or post-scrape browser actions. Requires a `scrape_id` from a prior scrape. + +## Search + +### Why use it + +Discover relevant pages from a query. Constrain to a site with `site:` in the query string (e.g. `site:docs.firecrawl.dev crawl webhooks`). + +### Preferred SDK method + +`client.search(query, **options)` → `SearchData` + +### Example + +```python +results = client.search( + "site:docs.firecrawl.dev webhook retries", + limit=5, + scrape_options={"formats": ["markdown"], "only_main_content": True}, +) + +for item in results.web or []: + print(getattr(item, "url", None), getattr(item, "title", None)) +``` + +**Wrong turn to avoid:** `search()` does not return `{ data: [...] }`. Access results via `result.web`, `result.news`, `result.images`, or `result.tools`. + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `query` | `str` | The search query. Use `site:example.com` to scope to a domain. | +| `sources` | `list[str \| Source]` | Which sources to search (`"web"`, `"news"`, `"images"`, `"alexandria"`). | +| `categories` | `list[str \| Category]` | Filter by category (`"github"`, `"research"`, `"pdf"`, `"developer"`). | +| `include_domains` | `list[str]` | Only include results from these domains. Mutually exclusive with `exclude_domains`. | +| `exclude_domains` | `list[str]` | Exclude results from these domains. Mutually exclusive with `include_domains`. | +| `limit` | `int` | Cap on number of results. Model default: `5`. | +| `tbs` | `str` | Time-based filter (e.g. `qdr:d` for past day, `qdr:w` for past week). | +| `location` | `str` | Location string for localized results. | +| `country` | `str` | ISO 3166-1 alpha-2 country code for geo-targeting. | +| `ignore_invalid_urls` | `bool` | Drop URLs that cannot be scraped. | +| `timeout` | `int` | Request timeout in milliseconds. Model default: `300000`. | +| `highlights` | `bool` | Generate query-relevant highlights. | +| `scrape_options` | `ScrapeOptions \| dict` | Scrape each search result with these options (see Scrape parameters). | +| `domain_tools` | `bool` | Include domain-matched Alexandria tools in results. | +| `tool_detail` | `"compact" \| "summary" \| "full"` | Detail level for discovered tools. | +| `enterprise` | `list[str]` | Enterprise options for zero data retention (`"zdr"`, `"anon"`). | +| `threat_protection` | `ThreatProtectionOptions` | Enterprise threat protection settings. | +| `integration` | `str` | Integration identifier for server-side tracking. | + +## Scrape + +### Why use it + +Get structured content from a URL in one or more formats — markdown, HTML, JSON extraction, screenshots, and more. + +### Preferred SDK method + +`client.scrape(url, **options)` → `Document` + +### Example + +```python +doc = client.scrape( + "https://example.com/pricing", + formats=[ + "markdown", + {"type": "json", "prompt": "Extract plan names and prices."}, + ], + only_main_content=True, +) + +print(doc.markdown) +print(doc.json) +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `url` | `str` | The URL to scrape. | +| `formats` | `list[FormatOption]` | Output formats. Strings: `"markdown"`, `"html"`, `"rawHtml"` (or `"raw_html"`), `"links"`, `"images"`, `"screenshot"`, `"summary"`, `"changeTracking"` (or `"change_tracking"`), `"attributes"`, `"branding"`, `"product"`, `"menu"`, `"audio"`, `"video"`. Objects: `{"type": "json", "prompt": ..., "schema": ...}`, `{"type": "question", "question": ...}`, `{"type": "highlights", "query": ...}`, `{"type": "screenshot", "fullPage": ..., "quality": ..., "viewport": ...}`, `{"type": "changeTracking", "modes": [...], "schema": ..., "prompt": ..., "tag": ...}`, `{"type": "attributes", "selectors": [{"selector": ..., "attribute": ...}]}`. | +| `headers` | `dict[str, str]` | Custom request headers. | +| `include_tags` | `list[str]` | Only include content from these HTML tags. | +| `exclude_tags` | `list[str]` | Exclude content from these HTML tags. | +| `only_main_content` | `bool` | Strip nav, footer, and boilerplate. | +| `timeout` | `int` | Timeout in milliseconds. | +| `wait_for` | `int` | Wait for the page to render (milliseconds). | +| `mobile` | `bool` | Emulate a mobile device. | +| `parsers` | `list[str \| PDFParser \| ImageParser]` | File parsing controls (e.g. `"pdf"` or `{"type": "pdf", "mode": "auto", "maxPages": 5}`). | +| `actions` | `list[Action]` | Pre-scrape browser actions. Types: `wait`, `click`, `write`, `press`, `scroll`, `screenshot`, `scrape`, `executeJavascript`, `pdf`. | +| `location` | `Location` | Geo and language-aware scraping (`{"country": "US", "languages": ["en-US"]}`). | +| `skip_tls_verification` | `bool` | Skip TLS verification. | +| `remove_base64_images` | `bool` | Drop base64 images from markdown. | +| `fast_mode` | `bool` | Faster scrapes with reduced fidelity. | +| `block_ads` | `bool` | Block ads and cookie popups. | +| `proxy` | `str` | Proxy control: `"basic"`, `"stealth"`, `"enhanced"`, `"auto"`. | +| `max_age` | `int` | Use cached data up to this age (milliseconds). | +| `store_in_cache` | `bool` | Cache the result in Firecrawl. | +| `lockdown` | `bool` | Only serve cached results, never make outbound requests. | +| `threat_protection` | `ThreatProtectionOptions` | Enterprise threat protection. | +| `profile` | `dict` | Persistent browser profile (`{"name": "...", "saveChanges": True}`). | +| `audit_metadata` | `AuditMetadata` | User attribution for SIEM logging. | +| `domain_tools` | `bool` | Include domain-matched Alexandria tools. | +| `tool_detail` | `"compact" \| "summary" \| "full"` | Detail level for discovered tools. | +| `integration` | `str` | Integration identifier. | + +## Interact + +### Why use it + +Control the browser session tied to a scrape job — run code or give natural-language instructions. Requires a `scrape_id` from a prior scrape's `metadata`. + +### Preferred SDK method + +`client.interact(job_id, code=None, *, prompt=None, language="node", timeout=None)` → `BrowserExecuteResponse` + +### Example + +```python +doc = client.scrape("https://example.com", formats=["markdown"]) +job_id = doc.metadata.get("scrapeId") + +result = client.interact( + job_id, + prompt="Click the pricing tab and summarize the plans.", +) +print(result.output) + +# Or run code directly: +code_result = client.interact( + job_id, + code="console.log(await page.title());", + language="node", + timeout=60, +) +print(code_result.stdout) + +# Stop the session when done: +client.stop_interaction(job_id) +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `job_id` | `str` | Scrape job ID from `document.metadata["scrapeId"]`. | +| `code` | `str` | Code to execute in the browser session. At least one of `code` or `prompt` required. | +| `prompt` | `str` | Natural-language instruction for the browser agent. At least one of `code` or `prompt` required. | +| `language` | `"python" \| "node" \| "bash"` | Runtime for code execution. Default: `"node"`. | +| `timeout` | `int` | Execution timeout in seconds (1–300). | +| `origin` | `str` | Optional origin label for telemetry. | + +Stop session: `client.stop_interaction(job_id)` → `BrowserDeleteResponse` + +## Notes + +- All parameter names use **snake_case**. Format strings accept both camelCase (`"rawHtml"`) and snake_case (`"raw_html"`). +- Deprecated aliases: `scrape_url` → `scrape`, `scrape_execute` → `interact`, `stop_interactive_browser` / `delete_scrape_browser` → `stop_interaction`. Always use the preferred names. +- `FirecrawlApp` is an alias for `Firecrawl` (backward compatibility). +- `SearchData` properties: `.web`, `.news`, `.images`, `.tools`. Accessing `.data` raises `AttributeError` with guidance. + +## Source Of Truth + +- `firecrawl/apps/python-sdk/firecrawl/client.py` +- `firecrawl/apps/python-sdk/firecrawl/v2/client.py` +- `firecrawl/apps/python-sdk/firecrawl/v2/types.py` +- `firecrawl-docs/api-reference/v2-openapi.json` diff --git a/agent-quickstart/rust.mdx b/agent-quickstart/rust.mdx new file mode 100644 index 000000000..a0759592a --- /dev/null +++ b/agent-quickstart/rust.mdx @@ -0,0 +1,245 @@ +--- +title: "Rust Agent Quickstart" +description: "Canonical Firecrawl Rust quickstart for external agents using search, scrape, and interact." +--- + +# Firecrawl Rust Agent Quickstart + +Canonical quickstart for external agents. Generated from SDK source (`firecrawl` crate **v2.21.0**) and the v2 OpenAPI spec. + +## Install + +```toml +[dependencies] +firecrawl = "2.21.0" +``` + +Requires an async runtime (tokio). + +## Authenticate + +```rust +use firecrawl::Client; + +let client = Client::new("fc-your-api-key")?; + +// Self-hosted: +// let client = Client::new_selfhosted("http://localhost:3002", Some("fc-your-api-key"))?; +``` + +`Client::new(api_key)` connects to `https://api.firecrawl.dev`. `Client::new_selfhosted(api_url, api_key)` connects to a custom URL — `api_key` is `Option` and can be `None` for the keyless free tier. + +## When To Use What + +- **search** — start with a query and need discovery. Returns ranked results you can then scrape. +- **scrape** — you already have a URL and want page content in one or more formats. +- **interact** — the page needs clicks, forms, or post-scrape browser actions. Requires a `scrape_id` from a prior scrape. + +## Search + +### Why use it + +Discover relevant pages from a query. Constrain to a site with `site:` in the query string (e.g. `site:docs.firecrawl.dev crawl webhooks`). + +### Preferred SDK method + +`client.search(query, options)` → `Result` + +### Example + +```rust +use firecrawl::{Client, SearchOptions, ScrapeOptions, Format}; + +let results = client + .search("site:docs.firecrawl.dev webhook retries", SearchOptions { + limit: Some(5), + scrape_options: Some(ScrapeOptions { + formats: Some(vec![Format::Markdown]), + only_main_content: Some(true), + ..Default::default() + }), + ..Default::default() + }) + .await?; + +if let Some(web) = &results.data.web { + for item in web { + // item is SearchResultOrDocument::WebResult or ::Document + println!("{:?}", item); + } +} +``` + +**Wrong turn to avoid:** `search()` returns `SearchResponse { data: SearchData { web, news, images, tools } }`. Access results via `data.web`, `data.news`, `data.images`, or `data.tools`. + +### Parameters + +All fields on `SearchOptions` are `Option` and default to `None` (omitted from the request). Use `..Default::default()` to fill unset fields. + +| Parameter | Type | Description | +|---|---|---| +| `query` | `impl AsRef` | The search query. Use `site:example.com` to scope to a domain. | +| `sources` | `Vec` | Which sources to search: `Web`, `News`, `Images`, `Alexandria`. | +| `categories` | `Vec` | Filter by category: `Github`, `Research`, `Pdf`. | +| `include_domains` | `Vec` | Only include results from these domains. | +| `exclude_domains` | `Vec` | Exclude results from these domains. | +| `limit` | `u32` | Cap on number of results. | +| `tbs` | `String` | Time-based filter (e.g. `"qdr:d"`, `"qdr:w"`). | +| `location` | `String` | Location string for localized results. | +| `country` | `String` | ISO country code for geo-targeting. | +| `ignore_invalid_urls` | `bool` | Drop URLs that cannot be scraped. | +| `timeout` | `u32` | Request timeout in milliseconds. | +| `highlights` | `bool` | Generate query-relevant highlights. Default: `true`. | +| `scrape_options` | `ScrapeOptions` | Scrape each search result (see Scrape parameters). | +| `domain_tools` | `bool` | Include domain-matched Alexandria tools. | +| `tool_detail` | `ToolDetail` | Detail level: `Compact`, `Summary`, `Full`. | +| `integration` | `String` | Integration identifier. | + +## Scrape + +### Why use it + +Get structured content from a URL in one or more formats — markdown, HTML, JSON extraction, screenshots, and more. + +### Preferred SDK method + +`client.scrape(url, options)` → `Result` + +### Example + +```rust +use firecrawl::{Client, ScrapeOptions, Format, JsonOptions}; + +let doc = client + .scrape("https://example.com/pricing", ScrapeOptions { + formats: Some(vec![Format::Markdown, Format::Json]), + json_options: Some(JsonOptions { + prompt: Some("Extract plan names and prices.".to_string()), + ..Default::default() + }), + only_main_content: Some(true), + ..Default::default() + }) + .await?; + +println!("{}", doc.markdown.unwrap_or_default()); +``` + +### Parameters + +All fields on `ScrapeOptions` are `Option` and default to `None`. Use `..Default::default()` to fill unset fields. + +| Parameter | Type | Description | +|---|---|---| +| `url` | `impl AsRef` | The URL to scrape. | +| `formats` | `Vec` | Output formats: `Markdown`, `Html`, `RawHtml`, `Links`, `Images`, `Screenshot`, `Summary`, `ChangeTracking`, `Json`, `Attributes`, `Branding`, `Product`, `Menu`, `Audio`, `Video`, `Question(QuestionFormat)`, `Highlights(HighlightsFormat)`. | +| `headers` | `HashMap` | Custom request headers. | +| `include_tags` | `Vec` | Only include content from these HTML tags. | +| `exclude_tags` | `Vec` | Exclude content from these HTML tags. | +| `only_main_content` | `bool` | Strip nav, footer, and boilerplate. | +| `timeout` | `u32` | Timeout in milliseconds. | +| `wait_for` | `u32` | Wait for the page to render (milliseconds). | +| `mobile` | `bool` | Emulate a mobile device. | +| `parsers` | `Vec` | File parsing controls (e.g. `ParserConfig::Pdf { parser_type, max_pages }`). | +| `actions` | `Vec` | Pre-scrape browser actions: `Wait`, `Click`, `Write`, `Press`, `Scroll`, `Screenshot`, `Scrape`, `ExecuteJavascript`, `Pdf`. | +| `location` | `LocationConfig` | Geo-aware scraping. Fields: `country`, `languages`. | +| `skip_tls_verification` | `bool` | Skip TLS verification. | +| `remove_base64_images` | `bool` | Drop base64 images from markdown. | +| `fast_mode` | `bool` | Faster scrapes with reduced fidelity. | +| `block_ads` | `bool` | Block ads and cookie popups. | +| `proxy` | `ProxyType` | Proxy control: `Basic`, `Stealth`, `Enhanced`, `Auto`. | +| `max_age` | `u32` | Use cached data up to this age (seconds). | +| `min_age` | `u32` | Use cached data only if at least this old (seconds). | +| `store_in_cache` | `bool` | Cache the result in Firecrawl. | +| `lockdown` | `bool` | Only serve cached results, never make outbound requests. | +| `redact_pii` | `bool` | Redact personally identifiable information. | +| `profile` | `ProfileConfig` | Persistent browser profile. Fields: `name`, `save_changes`. | +| `audit_metadata` | `AuditMetadata` | User attribution for SIEM logging. Field: `username`. | +| `domain_tools` | `bool` | Include domain-matched Alexandria tools. | +| `tool_detail` | `ToolDetail` | Detail level: `Compact`, `Summary`, `Full`. | +| `json_options` | `JsonOptions` | JSON extraction config. Fields: `schema`, `system_prompt`, `prompt`. | +| `screenshot_options` | `ScreenshotOptions` | Screenshot config. Fields: `full_page`, `quality`, `viewport`. | +| `change_tracking_options` | `ChangeTrackingOptions` | Change tracking config. Fields: `modes` (`GitDiff`, `Json`), `schema`, `prompt`, `tag`. | +| `attribute_selectors` | `Vec` | Attribute extraction. Fields: `selector`, `attribute`. | +| `integration` | `String` | Integration identifier. | + +## Interact + +### Why use it + +Control the browser session tied to a scrape job — run code or give natural-language instructions. Requires a scrape job ID from a prior scrape. + +### Preferred SDK method + +`client.interact(job_id, options)` → `Result` + +### Example + +```rust +use firecrawl::{Client, ScrapeOptions, ScrapeExecuteOptions, ScrapeExecuteLanguage, Format}; + +let doc = client + .scrape("https://example.com", ScrapeOptions { + formats: Some(vec![Format::Markdown]), + ..Default::default() + }) + .await?; + +let job_id = doc.metadata + .as_ref() + .and_then(|m| m.get("scrapeId")) + .and_then(|v| v.as_str()) + .expect("Missing scrapeId"); + +// Natural-language instruction: +let result = client + .interact(job_id, ScrapeExecuteOptions { + prompt: Some("Click the pricing tab and summarize the plans.".to_string()), + ..Default::default() + }) + .await?; +println!("{}", result.output.unwrap_or_default()); + +// Or run code directly: +let code_result = client + .interact(job_id, ScrapeExecuteOptions { + code: Some("console.log(await page.title());".to_string()), + language: Some(ScrapeExecuteLanguage::Node), + timeout: Some(60), + ..Default::default() + }) + .await?; +println!("{}", code_result.stdout.unwrap_or_default()); + +// Stop the session when done: +client.stop_interaction(job_id).await?; +``` + +### Parameters + +| Parameter | Type | Description | +|---|---|---| +| `job_id` | `impl AsRef` | Scrape job ID from `document.metadata["scrapeId"]`. | +| `code` | `Option` | Code to execute in the browser session. At least one of `code` or `prompt` required. | +| `prompt` | `Option` | Natural-language instruction for the browser agent. At least one of `code` or `prompt` required. | +| `language` | `ScrapeExecuteLanguage` | Runtime: `Python`, `Node`, `Bash`. Default: `Node`. | +| `timeout` | `u32` | Execution timeout in seconds. | + +Stop session: `client.stop_interaction(job_id)` → `Result` + +## Notes + +- All methods are **async** — requires a tokio runtime. +- Use `..Default::default()` to fill unset `Option` fields on structs. +- Passing `None` for `options` in `scrape()` and `search()` uses server defaults — no wrapping in `Some()` needed thanks to `impl Into>`. +- Wire format is camelCase (handled by serde). Rust fields use snake_case. +- Deprecated aliases: `scrape_execute` → `interact`, `stop_interactive_browser` / `delete_scrape_browser` → `stop_interaction`. +- SDK auto-sets `origin` to `"rust-sdk@{version}"` on scrape, search, and interact if not explicitly provided. + +## Source Of Truth + +- `firecrawl/apps/rust-sdk/src/v2/client.rs` +- `firecrawl/apps/rust-sdk/src/v2/search.rs` +- `firecrawl/apps/rust-sdk/src/v2/scrape.rs` +- `firecrawl/apps/rust-sdk/Cargo.toml` +- `firecrawl-docs/api-reference/v2-openapi.json`