Skip to content

Alexandria routing copy: name the catalogue, keep search as the front door, make sources opt-out explicit - #424

Merged
developersdigest merged 3 commits into
mainfrom
alexandria-routing-copy
Sep 22, 2026
Merged

developersdigest merged 3 commits into
mainfrom
alexandria-routing-copy

Conversation

@rakshith48

@rakshith48 rakshith48 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #405. Copy-only changes to the Alexandria routing text, plus one stale test assertion.

What changed

  • Server instructions (FULL_PROFILE_INSTRUCTIONS): the sentence "For structured records, filterable listings, transcripts, or datasets, first check firecrawl_find_tools…" is replaced by a description of what Alexandria is (data providers and workflows across companies, people, jobs, finance and filings, public records, real estate, places, retail and prices, developer data, news, research), when a provider beats scraping a page (same fields across several entities, provenance, exact figures, many records), an ordering rule ("before scraping more than one page for the same fields, spend one free find_tools call"), and an explicit statement that passing sources (including ["news"]) excludes provider matches. Search stays the front door; web results that already answer the question are used as they are.
  • firecrawl_find_tools description: opens with what the catalogue covers, keeps "Prefer normal firecrawl_search for a data task", and adds the scrape-first trigger. Identifier and title unchanged.
  • firecrawl_search sources description: "Passing sources excludes Alexandria provider matches, including sources: ["news"]; omit it unless you specifically need web-only or news-only results."
  • ALEXANDRIA_INSTRUCTIONS (search tool description): same framing, same vertical list, same sources statement.
  • firecrawl_scrape description: one line pointing scrape-first tasks at a search with sources unset.
  • tests/mcp-smoke.test.mjs: the instructions regex still expected the pre-ca120bc wording ("Alexandria offers ready-made workflows and provider tools") and failed on the branch head; updated to the new opener. Suite: 106/106.

Why these and not a rename

From the AX harness runs on this branch (EXP-053, 364 traces, Claude Code and Codex, 20 real tester queries; findings in firecrawl/agent-experience experiments/EXP-053-alexandria-mcp*.findings.md):

  • Renaming firecrawl_find_tools to firecrawl_alexandria or firecrawl_data_providers did not move bare-prompt routing. Claude Code only reaches a deferred MCP tool by keyword search or exact select, and its keyword queries are always verb-shaped ("firecrawl scrape search"); Codex made zero Firecrawl calls on any bare prompt regardless of name. So the identifier stays.
  • The sources opt-out was the mechanical blocker: on the earlier build agents passed sources: ["web"] on 300/350 searches and never saw a tools block. The ca120bc default wording fixed this for Claude (omits on 37/50 calls, and the tools block converted in 7 of 8 traces that received it) but not for Codex (supplied sources on 291/291 calls, never omitted). This copy says what supplying the argument costs rather than what omitting it gives, and covers ["news"], which suppressed the fitting provider on every portfolio-news trace.
  • The biggest remaining leak on Claude is tasks that never call search at all and go straight to firecrawl_scrape on a known URL; the scrape description line is the only text that reaches them.
  • Naming the capability is what routes: a prompt that says "Alexandria data providers" converted 9/11 provider-fit queries in both lanes against 0–5/11 for every arm that did not, so the noun now sits in the instructions and both descriptions.

Not in this PR

Two ergonomics bugs that cost calls in five traces: firecrawl_find_tools rejects a fully-qualified capability id (capabilities: ["zillow-com/properties/rental_search"] → invalid_option) although browse returns ids in that form; and sources: ["alexandria"] with toolDetail: "full" returns 59–80k characters and overflows the client token limit.

🤖 Generated with Claude Code


Summary by cubic

Rewrites the Alexandria routing copy in the server instructions and the firecrawl_search, firecrawl_find_tools, and firecrawl_scrape descriptions, based on 364 agent harness traces. The copy names the catalogue and its verticals, states that passing sources without alexandria in it excludes provider matches (including sources: ["web"] or ["news"]), and points scrape-first tasks at one search with sources unset. Search stays the front door; the tool identifier is unchanged. Also fixes the smoke-test assertion the copy change left stale and bumps the version to 3.25.2.

Routing rationale

  • Renaming firecrawl_find_tools never moved bare-prompt routing: Claude Code only reaches deferred MCP tools by keyword search, and Codex made no Firecrawl calls regardless of name.
  • Supplying sources was the mechanical blocker — agents passed sources: ["web"] on most searches and never saw the tools block — so the copy says what supplying the argument costs rather than what omitting it gives.
  • Naming the capability is what routes: prompts naming "Alexandria data providers" converted 9/11 provider-fit queries against 0–5/11 for other wording.

Review changes

  • Catalogue and sources copy extracted into shared constants so the instructions and tool descriptions stay in sync.
  • Sources opt-out scoped to source lists that omit alexandria; mixed lists that include it stay valid.
  • Scrape-first hint scoped to authenticated sessions with Alexandria access; keyless sessions scrape directly.
  • Smoke test now pins the catalogue opener, the sources opt-out, and the free find_tools call rule.

Written for commit 4e06782. Summary will update on new commits.

Review in cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread src/index.ts Outdated
Comment thread tests/mcp-smoke.test.mjs
Comment thread src/alexandria.ts Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files (changes from recent commits).

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread src/alexandria.ts Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 2 files (changes from recent commits).

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Requires human review: Rewrites Alexandria routing copy and tool descriptions to make sources opt-out explicit and steer scrape-first tasks to search, plus a test update. It intentionally changes agent routing toward provider matches, a product/operational tradeoff needing human sign-off.

Re-trigger cubic

…ront door, make sources opt-out explicit

Server instructions, firecrawl_search and firecrawl_find_tools descriptions, and the
sources parameter description rewritten from what agents did in 364 AX traces: name
Alexandria and its verticals through shared constants, say when a provider beats a page,
state that passing sources without alexandria (including ["news"]) excludes provider
matches, and give scrape-first tasks a pointer to run one search with sources unset
(authenticated sessions only). Tool identifier unchanged. Smoke test pins the catalogue
opener, the sources rule and the scrape-first rule.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@rakshith48
rakshith48 force-pushed the alexandria-routing-copy branch from 7eb817e to 6b3055a Compare September 22, 2026 03:56
@rakshith48
rakshith48 changed the base branch from alexandria-mcp to main September 22, 2026 03:57
rakshith48 and others added 2 commits September 22, 2026 12:16
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 4 files

Confidence score: 5/5

  • In src/index.ts, the duplicated Alexandria mode paragraph makes the firecrawl_scrape description unnecessarily repetitive, but it does not affect runtime behavior; remove the repeated line to keep the description clear.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/index.ts">

<violation number="1" location="src/index.ts:2452">
P2: The added line repeats the existing "Alexandria mode: pass `alexandria` ..." paragraph, so the firecrawl_scrape description now contains that paragraph twice (lines 2452–2453, the pre-existing line additionally carries the feedbackTool pointer sentence). This duplicate is emitted on every tools/list to every full-profile client, doubling that paragraph's tokens and diluting the routing copy this PR is meant to clarify. Delete the added line; the pre-existing full paragraph below it already covers the Alexandria mode.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread src/index.ts
Returns the selected content formats and page metadata. Authenticated responses can include a \`metadata.scrapeId\` for optional scrape feedback.

On an authenticated session with Alexandria access, if you are about to scrape the same fields from several pages, first run \`firecrawl_search\` with \`sources\` unset (or \`firecrawl_find_tools\`): a matching Alexandria provider returns those fields as typed records in one call. Keyless sessions have no provider matches; scrape directly.
Alexandria mode: pass \`alexandria\` (one \`{provider, capability, options}\` object or an array of 1-10) instead of \`url\` to execute catalogued Alexandria capabilities found through \`firecrawl_search\` sources \`alexandria\` or \`firecrawl_find_tools\`. The optional requestId identifies one logical execution: reuse the returned ID for retries of the identical payload, never a new ID to bypass pending or uncertain execution. Each call may include version to pin a published workflow; omitting it uses latest. Only timeout also applies at the top level in this mode. Returns per-capability results in \`data.alexandria\`, including \`data\`, \`records\`, or an \`error\` with a code. Check each item for errors even when the outer response is successful. Alexandria needs an API key on a team with Alexandria enabled.

@cubic-dev-ai cubic-dev-ai Bot Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The added line repeats the existing "Alexandria mode: pass alexandria ..." paragraph, so the firecrawl_scrape description now contains that paragraph twice (lines 2452–2453, the pre-existing line additionally carries the feedbackTool pointer sentence). This duplicate is emitted on every tools/list to every full-profile client, doubling that paragraph's tokens and diluting the routing copy this PR is meant to clarify. Delete the added line; the pre-existing full paragraph below it already covers the Alexandria mode.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/index.ts, line 2452:

<comment>The added line repeats the existing "Alexandria mode: pass `alexandria` ..." paragraph, so the firecrawl_scrape description now contains that paragraph twice (lines 2452–2453, the pre-existing line additionally carries the feedbackTool pointer sentence). This duplicate is emitted on every tools/list to every full-profile client, doubling that paragraph's tokens and diluting the routing copy this PR is meant to clarify. Delete the added line; the pre-existing full paragraph below it already covers the Alexandria mode.</comment>

<file context>
@@ -2444,6 +2448,8 @@ Firecrawl may reuse recently indexed content instead of refetching the page, and
 Returns the selected content formats and page metadata. Authenticated responses can include a \`metadata.scrapeId\` for optional scrape feedback.
 
+On an authenticated session with Alexandria access, if you are about to scrape the same fields from several pages, first run \`firecrawl_search\` with \`sources\` unset (or \`firecrawl_find_tools\`): a matching Alexandria provider returns those fields as typed records in one call. Keyless sessions have no provider matches; scrape directly.
+Alexandria mode: pass \`alexandria\` (one \`{provider, capability, options}\` object or an array of 1-10) instead of \`url\` to execute catalogued Alexandria capabilities found through \`firecrawl_search\` sources \`alexandria\` or \`firecrawl_find_tools\`. The optional requestId identifies one logical execution: reuse the returned ID for retries of the identical payload, never a new ID to bypass pending or uncertain execution. Each call may include version to pin a published workflow; omitting it uses latest. Only timeout also applies at the top level in this mode. Returns per-capability results in \`data.alexandria\`, including \`data\`, \`records\`, or an \`error\` with a code. Check each item for errors even when the outer response is successful. Alexandria needs an API key on a team with Alexandria enabled.
 Alexandria mode: pass \`alexandria\` (one \`{provider, capability, options}\` object or an array of 1-10) instead of \`url\` to execute catalogued Alexandria capabilities found through \`firecrawl_search\` sources \`alexandria\` or \`firecrawl_find_tools\`. The optional requestId identifies one logical execution: reuse the returned ID for retries of the identical payload, never a new ID to bypass pending or uncertain execution. Each call may include version to pin a published workflow; omitting it uses latest. Only timeout also applies at the top level in this mode. Returns per-capability results in \`data.alexandria\`, including \`data\`, \`records\`, or an \`error\` with a code. Check each item for errors even when the outer response is successful. Alexandria needs an API key on a team with Alexandria enabled. Alexandria results include a \`feedbackTool\` pointer: after the task, report how the catalogue served the website through \`firecrawl_feedback\` with endpoint \`alexandria\` (free, no job ID).
 
</file context>
Fix with cubic

@developersdigest
developersdigest merged commit 1d0ca84 into main Sep 22, 2026
2 checks passed
rakshith48 added a commit that referenced this pull request Sep 22, 2026
The only conflict was the rewritten profile instructions; main's text is
kept with the threadId sentence re-appended.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants