Add Alexandria search, progressive discovery and scrape execution to MCP - #405
Conversation
- firecrawl_search: sources enum gains "exchange" in both profiles;
data.exchange[] and creditsUsed pass through untouched
- firecrawl_scrape: url becomes optional, exchange[1..10]{provider,
capability, options?} is accepted (exactly one of url or exchange, no
other scrape options alongside) and posts /v2/scrape {exchange, origin}
returning {success, scrape_id, data:{exchange, creditsCost}} as-is
- new firecrawl_exchange_discover tool: GET /exchange/discover
[/cohort[/provider[/capability]]] with q/limit on the index route only
and expand on cohort routes
- Exchange error bodies (403 no flag, 501 semantic_not_configured, future
402/409) are relayed in-band with their code; 401 keeps credential
recovery
- keyless and unauthenticated sessions get "Exchange requires an API key
on a team with Exchange access"; the discover tool stays out of
KEYLESS_TOOL_NAMES so hosted keyless sessions never list it
- instructions, tool descriptions, README tool list, CHANGELOG, version
3.25.0
- tests/mcp-exchange.test.mjs covers outbound bodies, discover path
building and validation, error relay, local and hosted keyless
rejection, and agent-metadata policy on the new language
Claude-Session: https://claude.ai/code/session_01KpbshKPNuCDBD4TLJ8T2UQ
…paths
Review follow-up for the Exchange v2 MCP surface:
- firecrawl_search `sources` entries may now be a bare source name
("web" | "images" | "news" | "exchange") or {type} in both profiles,
matching the frozen /v2/search contract and the M1 acceptance call
sources: ["exchange"]; entries are forwarded verbatim.
- firecrawl_exchange_discover refuses "." and ".." cohort, provider, and
capability segments and refuses `expand` off the cohort route before
any request, so the tool stays strictly under /exchange/discover*.
- relayExchangeError also surfaces `chargeId` from the reserved 409
duplicate_request / request_in_flight billing errors.
- Tests: bare-string sources forwarded verbatim (full and search
profiles) and an unknown source rejected without a request; 409
chargeId relay; dot-segment and expand-off-cohort refusals with zero
API calls; sources schema assertion updated to the union shape.
- docs/search-profile.md documents the exchange source on the search
surface; CHANGELOG updated.
Claude-Session: https://claude.ai/code/session_01KpbshKPNuCDBD4TLJ8T2UQ
Exchange capability addresses are provider-relative (e.g. series/observations for fred), not finance/series/observations. Fix README, tool schema descriptions and test fixtures; no behaviour change. Claude-Session: https://claude.ai/code/session_01KpbshKPNuCDBD4TLJ8T2UQ
Relay THIRD_PARTY_DATA_TERMS_REQUIRED with its requiresAction payload and next_actions from every Exchange call site, the plain scrape/search SDK paths, and keylessPost. The server never accepts terms; the tool result names the admin URL and tells the agent to retry the identical payload and requestId after a person confirms.
There was a problem hiding this comment.
All reported issues were addressed across 10 files
Tip: instead of fixing issues one by one fix them all with cubic
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 4 files (changes from recent commits).
Heads up: you’re close to your flex budget. Increase your flex budget so reviews don’t pause.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 4 files (changes from recent commits).
Heads up: you’re close to your flex budget. Increase your flex budget so reviews don’t pause.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 7 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Heads up: you’re close to your flex budget. Increase your flex budget so reviews don’t pause.
Requires human review: Adds Alexandria discovery/execution tools and a new CLI credential-loading path. Expands the API surface and introduces a new authentication mechanism; terms acceptance depends on an external endpoint not visible in this diff.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 5 files (changes from recent commits).
Heads up: you’re close to your flex budget. Increase your flex budget so reviews don’t pause.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 4 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 3 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 3 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Requires human review: Auto-approval blocked by 5 unresolved issues from previous reviews.
Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 6 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Requires human review: Auto-approval blocked by 5 unresolved issues from previous reviews.
Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 1 file (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Requires human review: Auto-approval blocked by 5 unresolved issues from previous reviews.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 2 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 4 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Fix all with cubic | Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 4 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Requires human review: Auto-approval blocked by 1 unresolved issue from previous reviews.
Re-trigger cubic
Adds the missing_capability capability issue (provider exists but lacks the capability). Unlike new_capability_request it does not require requestedFunctionality. Tool description and README list the capability issue codes; smoke test covers acceptance without requestedFunctionality and rejection of unknown issue codes. Co-authored-by: Cursor <cursoragent@cursor.com>
|
c9d2f14 adds the new Server-side, Alexandria feedback storage is moving to dedicated tables (firecrawl-db PR https://github.com/firecrawl/firecrawl-db/pull/301). There is no client-visible contract change beyond the new enum value. Local run: |
There was a problem hiding this comment.
0 issues found across 4 files (changes from recent commits).
Confidence score: 5/5
- Automated review surfaced no issues in the provided summaries.
- No files require special attention.
Requires human review: Auto-approval blocked by 1 unresolved issue from previous reviews.
Re-trigger cubic
…ry-routing Three conflicts, all where Alexandria (#405) and the query router touch the same lines. Both firecrawl_search surfaces: main added an Alexandria credential guard at exactly the point this branch inserts the routing call. Kept both, guard first — it is a cheap local check that fails fast, where routing makes a network call. `exchangeSource` is still read further down, so its declaration stays. CHANGELOG: main renamed [Unreleased] to [3.25.0], so the router entry folds into that section rather than leaving two release headings. One interaction the merge creates rather than inherits: `domainTools` now blocks the research retarget, the same way `scrapeOptions` already does. Alexandria tool suggestions ride along with web results and the paper index has none to offer, so retargeting would silently drop what the caller asked for. New reason token `domain_tools_requested`, with a test. 133 tests pass. The three eslint errors in src/alexandria.ts are byte-identical to main and predate this merge; CI runs `pnpm test`, not lint. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Follow-up: agents using Alexandria purely through MCP were never told to send feedback (the |
Resolves conflicts in the profile instructions, the scrape tool schema, the feedback schema and the CHANGELOG, and threads the new Alexandria surface: firecrawl_find_tools declares threadId (its schema is strict), the Alexandria-mode refinement on firecrawl_scrape allows threadId alongside alexandria, and result stamping keeps compact JSON compact now that main emits compact results. Two Alexandria tests that deep-compare payloads split off the thread ID first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Adds structured tool discovery and execution to Search and Scrape, with progressive catalogue inspection, provider terms, optional version selection, and guidance for retained results.
Search defaults to compact tool identities and descriptions; explicit summary/full detail remains available. Selected contract inspection prefers inputs and output shape, requesting examples when needed. URL Scrape keeps its summary default. Large-result guidance preserves request identity and combines selective reads through retained-result Bash.
No new credential-file loading or deployment environment variables. Existing shared authentication recovery and unrelated tool formatting remain unchanged. This does not automatically intercept client-side context overflow.
Validation: all 110 local tests passed during scope review; subsequent instruction updates passed 10 focused discovery/scrape tests and build. Live CLI/MCP trials covered discovery, contract inspection, execution, and retained-result analysis. An eight-task instruction comparison completed seven tasks; one source execution returned an unresolved error. Selective contracts reduced discovery output versus full expansion, and successful large-result trials used bounded remote projections.