Skip to content

[Feature]: Buddy search parity: source/date filters, grep tool and streamed reasoning #206

Description

@daniilperkin

Parent Story

Part of SprintStartProject/Wiki#319 (retire /chat, Buddy becomes the only conversation surface). Blocks SprintStartProject/sprintstart-backend#214.

Order: start after open PR sprintstart-ai#208 (DavidLeuter, Buddy as onboarding tutor, opened 2026-09-25) is merged. It changes buddy_agent.py and adds 268 lines to buddy_persona.py, the same files items 4 and 7–9 touch.

Summary & Goal

Give the Buddy agent everything the chat agent has for answering questions, so capabilities-off Buddy can replace /chat:

  1. Source-system and date filters on retrieval.
  2. The grep tool next to search_docs.
  3. The model's reasoning, returned with each turn, so the reasoning/tool panel moves over.

What already exists on dev (verified 2026-09-24)

  • RetrievalFilters (src/rag/types.py) already carries source_systems, time_from, time_to next to project_ids, and rag/filters.py applies them on both halves of hybrid retrieval.
  • ChatFilters (src/api/schemas.py) already validates the request shape and normalises source systems to upper case. routes/chat.py::_retrieval_filters_from_request turns it into RetrievalFilters.
  • ChatAgent (src/agents/chat_agent.py) mounts RetrieveTool + GrepTool (agents/tools/grep.py) with the same filters and yields ReasoningDelta as Reasoning events.
  • buddy_agent.run_agent_turn (src/onboarding/buddy_agent.py) only has _SEARCH_TOOL (query only), calls retrieve(..., top_k=5, filters=RetrievalFilters(project_ids=project_ids)), and emits no reasoning. capabilities_enabled already reaches it from BuddyAgentRequest.

Technical Specification

  1. Request. Add filters: ChatFilters | None = None to BuddyAgentRequest. Reuse ChatFilters (rename it to a neutral RetrievalFiltersSchema only if the chat route keeps working unchanged); do not add a second schema with the same fields.
  2. Filters are the hire's, not the model's. Build one RetrievalFilters(project_ids=..., source_systems=..., time_from=..., time_to=...) per turn and pass it to every retrieval call. Do not add filter arguments to the search_docs tool schema: the model must not be able to widen what the hire narrowed.
  3. Project scope is unchanged. Backend decides it: a project-scoped conversation sends one project id, an unscoped one sends all of the hire's projects (Wiki#319, decision 1). Keep drop_test_material and the fail-closed project rule from AGENTS.md.
  4. grep when capabilities are off. Mount a grep tool spec next to search_docs and execute it with GrepTool(store, exclusions=..., filters=...). It scans all_chunks_without_embeddings(), so it must get the same filters (see AGENTS.md, "Project separation"). Deduplicate citations across both tools by chunk id, as search_docs does today.
  5. Reasoning (batch, decided 2026-09-25). POST /api/v1/onboarding/buddy/agent stays a synchronous JSON endpoint (the final / pending_tool_calls loop), so there is nothing to stream. Instead add reasoning: list[str] to BuddyAgentResponse: one entry per model step of this call, taken from ChatResult.reasoning, which llm.chat() already returns. It is never part of text or messages. The backend emits it as reasoning events before the answer tokens (backend#214), so it shows after the model has thought, not live.
  6. Empty filtered result. When filters are set and retrieval returns nothing, answer with chat's canned "no matching sources for the selected filters" reply instead of an ungrounded answer (same rule as _has_narrowing_filters).
  7. Prompt-injection fence (chat has it, Buddy does not). ChatAgent wraps the user's question in a random marker (wrap_user_query) and adds _QUERY_FENCE_NOTE, so text in the question cannot act as instructions. run_agent_turn sends the question unwrapped. Apply the same fence to the Buddy turn in both modes. Keep the wrapped bytes stable across hops, because of the prompt cache.
  8. Evidence budget parity. Chat returns up to _MAX_EVIDENCE_CHUNKS = 12 / _MAX_EVIDENCE_CHARS = 8_000 per step and runs up to 4 tool calls in parallel. Buddy's search_docs takes _TOP_K = 5 with _MIN_SCORE = 0.3. For capabilities off, match chat's budget; move chat's evidence-budget selection out of chat_agent.py into a shared module and use it from both, rather than copying it: ai#207 deletes chat_agent.py.
  9. Search-only prompt. _SEARCH_ONLY_CLAUSE (buddy_persona.py) says "you have search_docs and nothing else", which is wrong once grep is mounted. Add chat's guidance: prefer search_docs for concepts and grep for exact identifiers; request several searches in one turn; answer only from the results and say so plainly when they fall short; do not search for greetings or small talk. Also add _FINAL_ANSWER_INSTRUCTION's rule that source text is untrusted data.

Note on 5: live reasoning and true token streaming would mean turning the agent endpoint into SSE and reworking the backend's tool loop (OnboardingAiClient.buddyAgentTurn, BuddyService, BuddyReplyStream). That is out of scope here; the batch field keeps llm.chat() and the prompt-cache byte layout unchanged.

Out of Scope

  • Retiring /api/v1/chat, ChatAgent, ChatOrchestrator (sprintstart-ai#207).
  • Live reasoning and true token streaming (SSE agent endpoint).
  • grep with capabilities on, and team mode.

Acceptance Criteria

  • BuddyAgentRequest accepts optional filters; omitting them keeps today's behaviour
  • search_docs and grep both apply project, source-system and time filters
  • The model cannot change the filters through tool arguments
  • Capabilities off mounts search_docs + grep and no backend action tools
  • BuddyAgentResponse.reasoning carries each step's reasoning (empty list when the model gives none) and never lands in text or messages
  • Filters that match nothing give the canned no-sources reply
  • The question is fenced with a random marker in the Buddy turn; an injected "ignore your rules" in the question is treated as data
  • Capabilities-off evidence budget matches chat's (12 chunks / 8,000 chars per step)
  • _SEARCH_ONLY_CLAUSE names both tools and carries chat's grounding rules
  • Unit tests (with ScriptedLLMClient / StubVectorStore): filtered retrieval, grep filtering, no-widening, reasoning field, empty filtered result
  • Gates pass: uv run ruff format --check ., uv run ruff check ., uv run pyright src/, uv run pytest

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions