Parent Story
Part of SprintStartProject/Wiki#319 (retire /chat, Buddy becomes the only conversation surface). Blocks SprintStartProject/sprintstart-backend#214.
Order: start after open PR sprintstart-ai#208 (DavidLeuter, Buddy as onboarding tutor, opened 2026-09-25) is merged. It changes buddy_agent.py and adds 268 lines to buddy_persona.py, the same files items 4 and 7–9 touch.
Summary & Goal
Give the Buddy agent everything the chat agent has for answering questions, so capabilities-off Buddy can replace /chat:
- Source-system and date filters on retrieval.
- The
grep tool next to search_docs.
- The model's reasoning, returned with each turn, so the reasoning/tool panel moves over.
What already exists on dev (verified 2026-09-24)
RetrievalFilters (src/rag/types.py) already carries source_systems, time_from, time_to next to project_ids, and rag/filters.py applies them on both halves of hybrid retrieval.
ChatFilters (src/api/schemas.py) already validates the request shape and normalises source systems to upper case. routes/chat.py::_retrieval_filters_from_request turns it into RetrievalFilters.
ChatAgent (src/agents/chat_agent.py) mounts RetrieveTool + GrepTool (agents/tools/grep.py) with the same filters and yields ReasoningDelta as Reasoning events.
buddy_agent.run_agent_turn (src/onboarding/buddy_agent.py) only has _SEARCH_TOOL (query only), calls retrieve(..., top_k=5, filters=RetrievalFilters(project_ids=project_ids)), and emits no reasoning. capabilities_enabled already reaches it from BuddyAgentRequest.
Technical Specification
- Request. Add
filters: ChatFilters | None = None to BuddyAgentRequest. Reuse ChatFilters (rename it to a neutral RetrievalFiltersSchema only if the chat route keeps working unchanged); do not add a second schema with the same fields.
- Filters are the hire's, not the model's. Build one
RetrievalFilters(project_ids=..., source_systems=..., time_from=..., time_to=...) per turn and pass it to every retrieval call. Do not add filter arguments to the search_docs tool schema: the model must not be able to widen what the hire narrowed.
- Project scope is unchanged. Backend decides it: a project-scoped conversation sends one project id, an unscoped one sends all of the hire's projects (Wiki#319, decision 1). Keep
drop_test_material and the fail-closed project rule from AGENTS.md.
grep when capabilities are off. Mount a grep tool spec next to search_docs and execute it with GrepTool(store, exclusions=..., filters=...). It scans all_chunks_without_embeddings(), so it must get the same filters (see AGENTS.md, "Project separation"). Deduplicate citations across both tools by chunk id, as search_docs does today.
- Reasoning (batch, decided 2026-09-25).
POST /api/v1/onboarding/buddy/agent stays a synchronous JSON endpoint (the final / pending_tool_calls loop), so there is nothing to stream. Instead add reasoning: list[str] to BuddyAgentResponse: one entry per model step of this call, taken from ChatResult.reasoning, which llm.chat() already returns. It is never part of text or messages. The backend emits it as reasoning events before the answer tokens (backend#214), so it shows after the model has thought, not live.
- Empty filtered result. When filters are set and retrieval returns nothing, answer with chat's canned "no matching sources for the selected filters" reply instead of an ungrounded answer (same rule as
_has_narrowing_filters).
- Prompt-injection fence (chat has it, Buddy does not).
ChatAgent wraps the user's question in a random marker (wrap_user_query) and adds _QUERY_FENCE_NOTE, so text in the question cannot act as instructions. run_agent_turn sends the question unwrapped. Apply the same fence to the Buddy turn in both modes. Keep the wrapped bytes stable across hops, because of the prompt cache.
- Evidence budget parity. Chat returns up to
_MAX_EVIDENCE_CHUNKS = 12 / _MAX_EVIDENCE_CHARS = 8_000 per step and runs up to 4 tool calls in parallel. Buddy's search_docs takes _TOP_K = 5 with _MIN_SCORE = 0.3. For capabilities off, match chat's budget; move chat's evidence-budget selection out of chat_agent.py into a shared module and use it from both, rather than copying it: ai#207 deletes chat_agent.py.
- Search-only prompt.
_SEARCH_ONLY_CLAUSE (buddy_persona.py) says "you have search_docs and nothing else", which is wrong once grep is mounted. Add chat's guidance: prefer search_docs for concepts and grep for exact identifiers; request several searches in one turn; answer only from the results and say so plainly when they fall short; do not search for greetings or small talk. Also add _FINAL_ANSWER_INSTRUCTION's rule that source text is untrusted data.
Note on 5: live reasoning and true token streaming would mean turning the agent endpoint into SSE and reworking the backend's tool loop (OnboardingAiClient.buddyAgentTurn, BuddyService, BuddyReplyStream). That is out of scope here; the batch field keeps llm.chat() and the prompt-cache byte layout unchanged.
Out of Scope
- Retiring
/api/v1/chat, ChatAgent, ChatOrchestrator (sprintstart-ai#207).
- Live reasoning and true token streaming (SSE agent endpoint).
grep with capabilities on, and team mode.
Acceptance Criteria
Parent Story
Part of SprintStartProject/Wiki#319 (retire
/chat, Buddy becomes the only conversation surface). Blocks SprintStartProject/sprintstart-backend#214.Order: start after open PR sprintstart-ai#208 (DavidLeuter, Buddy as onboarding tutor, opened 2026-09-25) is merged. It changes
buddy_agent.pyand adds 268 lines tobuddy_persona.py, the same files items 4 and 7–9 touch.Summary & Goal
Give the Buddy agent everything the chat agent has for answering questions, so capabilities-off Buddy can replace
/chat:greptool next tosearch_docs.What already exists on
dev(verified 2026-09-24)RetrievalFilters(src/rag/types.py) already carriessource_systems,time_from,time_tonext toproject_ids, andrag/filters.pyapplies them on both halves of hybrid retrieval.ChatFilters(src/api/schemas.py) already validates the request shape and normalises source systems to upper case.routes/chat.py::_retrieval_filters_from_requestturns it intoRetrievalFilters.ChatAgent(src/agents/chat_agent.py) mountsRetrieveTool+GrepTool(agents/tools/grep.py) with the same filters and yieldsReasoningDeltaasReasoningevents.buddy_agent.run_agent_turn(src/onboarding/buddy_agent.py) only has_SEARCH_TOOL(queryonly), callsretrieve(..., top_k=5, filters=RetrievalFilters(project_ids=project_ids)), and emits no reasoning.capabilities_enabledalready reaches it fromBuddyAgentRequest.Technical Specification
filters: ChatFilters | None = NonetoBuddyAgentRequest. ReuseChatFilters(rename it to a neutralRetrievalFiltersSchemaonly if the chat route keeps working unchanged); do not add a second schema with the same fields.RetrievalFilters(project_ids=..., source_systems=..., time_from=..., time_to=...)per turn and pass it to every retrieval call. Do not add filter arguments to thesearch_docstool schema: the model must not be able to widen what the hire narrowed.drop_test_materialand the fail-closed project rule fromAGENTS.md.grepwhen capabilities are off. Mount agreptool spec next tosearch_docsand execute it withGrepTool(store, exclusions=..., filters=...). It scansall_chunks_without_embeddings(), so it must get the same filters (seeAGENTS.md, "Project separation"). Deduplicate citations across both tools by chunk id, assearch_docsdoes today.POST /api/v1/onboarding/buddy/agentstays a synchronous JSON endpoint (thefinal/pending_tool_callsloop), so there is nothing to stream. Instead addreasoning: list[str]toBuddyAgentResponse: one entry per model step of this call, taken fromChatResult.reasoning, whichllm.chat()already returns. It is never part oftextormessages. The backend emits it asreasoningevents before the answer tokens (backend#214), so it shows after the model has thought, not live._has_narrowing_filters).ChatAgentwraps the user's question in a random marker (wrap_user_query) and adds_QUERY_FENCE_NOTE, so text in the question cannot act as instructions.run_agent_turnsends the question unwrapped. Apply the same fence to the Buddy turn in both modes. Keep the wrapped bytes stable across hops, because of the prompt cache._MAX_EVIDENCE_CHUNKS = 12/_MAX_EVIDENCE_CHARS = 8_000per step and runs up to 4 tool calls in parallel. Buddy'ssearch_docstakes_TOP_K = 5with_MIN_SCORE = 0.3. For capabilities off, match chat's budget; move chat's evidence-budget selection out ofchat_agent.pyinto a shared module and use it from both, rather than copying it: ai#207 deleteschat_agent.py._SEARCH_ONLY_CLAUSE(buddy_persona.py) says "you havesearch_docsand nothing else", which is wrong oncegrepis mounted. Add chat's guidance: prefersearch_docsfor concepts andgrepfor exact identifiers; request several searches in one turn; answer only from the results and say so plainly when they fall short; do not search for greetings or small talk. Also add_FINAL_ANSWER_INSTRUCTION's rule that source text is untrusted data.Note on 5: live reasoning and true token streaming would mean turning the agent endpoint into SSE and reworking the backend's tool loop (
OnboardingAiClient.buddyAgentTurn,BuddyService,BuddyReplyStream). That is out of scope here; the batch field keepsllm.chat()and the prompt-cache byte layout unchanged.Out of Scope
/api/v1/chat,ChatAgent,ChatOrchestrator(sprintstart-ai#207).grepwith capabilities on, and team mode.Acceptance Criteria
BuddyAgentRequestaccepts optionalfilters; omitting them keeps today's behavioursearch_docsandgrepboth apply project, source-system and time filterssearch_docs+grepand no backend action toolsBuddyAgentResponse.reasoningcarries each step's reasoning (empty list when the model gives none) and never lands intextormessages_SEARCH_ONLY_CLAUSEnames both tools and carries chat's grounding rulesScriptedLLMClient/StubVectorStore): filtered retrieval, grep filtering, no-widening, reasoning field, empty filtered resultuv run ruff format --check .,uv run ruff check .,uv run pyright src/,uv run pytest