Skip to content

Add file-level query mode for full-file retrieval #181

Description

@Lucas-Bur

Problem Statement

pix query currently returns ranked source chunks. That is useful for precise retrieval, but it is fragmented for humans and for agents that need a complete file as the next unit of context. A caller can discover a relevant chunk without receiving the complete source file that should be inspected or passed on.

Solution

Add a second query mode, file, alongside the existing chunk mode. Retrieval remains chunk-based internally, but the already-fused chunk results are deduplicated and ranked as unique files before top is applied. A file result returns the complete text of the source file, or only file metadata when noContent is requested.

The CLI exposes the mode as --file. The transport-independent query contract, MCP tools, and persisted aliases expose an explicit mode value. The existing chunk mode remains the default for a direct query without --file.

User Stories

  1. As a developer, I want to search for a concept and receive complete matching files, so that I can understand the surrounding implementation without manually opening every chunk.
  2. As an AI agent, I want a file-level query mode, so that I can pass a relevant source file into a later reasoning step as one coherent context unit.
  3. As a developer, I want chunk-level retrieval to remain the default, so that existing precise searches keep their current behavior.
  4. As a developer, I want --file on pix query, so that I can select file-level retrieval without learning a second command.
  5. As an MCP client, I want an explicit query mode, so that I can request chunk or file results through the same shared query API.
  6. As a developer, I want top to mean the number of unique files in file mode, so that one file with several matching chunks does not consume several result slots.
  7. As a developer, I want a file to appear at most once, so that the result list is not duplicated by multiple matching chunks.
  8. As a developer, I want file ranking to preserve the strongest matching chunk, so that a file with one excellent match is not penalized for lacking additional matches.
  9. As a developer, I want deterministic tie-breaking between equally ranked files, so that repeated queries return stable ordering.
  10. As a developer, I want the complete textual source of each selected file, so that file mode provides a meaningful alternative to chunk mode rather than a grouped list of chunks.
  11. As a developer, I want noContent to work in file mode, so that I can discover file paths and relevance metadata without loading source text.
  12. As a developer, I want chunk line ranges and context fields omitted from file results, so that consumers cannot mistake a file result for a chunk result.
  13. As a developer, I want contextLines to have no effect in file mode, so that the full-file contract remains unambiguous.
  14. As a developer, I want maxCharacters to protect file-mode output, so that a large file cannot unexpectedly exhaust the response budget.
  15. As a developer, I want truncated output to be explicitly marked, so that agents can distinguish a complete file from a budget-limited file.
  16. As a developer, I want file aliases to preserve their file mode, so that a saved shortcut consistently produces file results.
  17. As a developer, I do not want a file alias to silently switch back to chunk mode, so that aliases remain stable semantic shortcuts.
  18. As a developer, I want clipboard copy to retain the existing formatted-text behavior, so that adding file mode does not unexpectedly change an established output path.
  19. As a maintainer, I want file mode to reuse the existing chunk retrieval and source-loading boundaries, so that the feature does not introduce a second embedding or indexing system.
  20. As a maintainer, I want the feature to work consistently in CLI, aliases, and MCP, so that each transport exposes the same retrieval semantics.

Implementation Decisions

  • Add two explicit query modes: chunk and file. chunk preserves the current result semantics and remains the direct-query default; file is selected by the CLI --file flag or the corresponding explicit API/MCP mode.
  • Keep all existing retrieval channels, evidence routing, and chunk-level fusion unchanged. File mode is a result-level transformation, not a new retrieval channel and not a file-level embedding/index.
  • Aggregate after chunk-level fusion and path filtering, but before the result limit and source hydration. The aggregator consumes the complete available fused result ordering rather than treating top as a chunk candidate limit.
  • Group fused chunk results by normalized repository-relative file path. Each file gets the score and relevance of its strongest fused chunk (max). This is the MVP metric because it preserves a strong single hit and avoids the length bias of unbounded summation.
  • Preserve the existing fused ordering as the first tie-breaker: the file whose strongest chunk occurs first wins. Use the normalized file path as the final deterministic tie-breaker.
  • Apply top after file aggregation. In file mode, top means the number of unique file results returned; the existing top-range clamping remains in effect.
  • A file result is a distinct result shape containing file, score, rel, and optional text. It does not expose chunk line ranges, chunk context, match spans, or individual chunks.
  • Hydrate the complete textual source for selected files through the existing source-loading boundary, preserving the same textual source from which indexed chunks were produced. Do not reconstruct a file by concatenating retrieved chunks.
  • noContent omits text from file results and returns file metadata only. It also means there is no content to truncate.
  • contextLines has no effect in file mode because the selected content is already the complete file.
  • maxCharacters remains output logic after retrieval, result selection, and hydration. It does not change chunking, embedding, scoring, or ranking.
  • The existing chunk-mode character-budget behavior remains intact: the budget covers rendered metadata, text, and context; the last possible result may be shortened with the existing [...] marker, context is removed when that happens, and later results are not emitted.
  • File mode applies the same output-budget principle to file metadata plus complete file text. A last file may be shortened to fit the remaining budget.
  • Add an optional machine-readable truncated: true marker whenever the character budget actually changes a result's text. This applies consistently to chunk and file results; the field is omitted for complete results and for noContent results.
  • Preserve the existing formatted-text clipboard contract. File-mode copy contains the formatted file result and complete file text, with the same result separators as the existing copy path; it does not switch clipboard output to JSON.
  • Persist mode in query aliases. pix alias add ... --file creates a file alias; alias execution has no mode override, so a file alias remains a file alias and a chunk alias remains a chunk alias.
  • Newly written alias entries must contain an explicit mode. Existing alias files that omit the required mode are allowed to fail schema validation and must be updated; no compatibility migration is required.
  • The query response identifies its mode and exposes the corresponding chunk-result or file-result shape. CLI JSON, human output, aliases, and MCP use the same underlying semantics.
  • No hidden maximum file-size limit is introduced. Users control source loading with noContent and output protection with maxCharacters.
  • A file result can only originate from at least one ranked chunk. Empty files are not independently ranked or added to the index by this feature.
  • A bounded multi-hit or dis_max-style bonus is explicitly deferred; it may be evaluated later against the retrieval benchmark and length-bias fixtures.

Acceptance Criteria

  • Direct queries support chunk mode and file mode, with chunk mode as the default and --file selecting file mode.
  • The shared query and MCP contracts represent the selected mode explicitly.
  • File-mode results contain unique files only and never expose chunk line ranges, context fields, match spans, or individual chunks.
  • File-mode file scores use the strongest fused chunk score and relevance for that file.
  • File-mode top is applied after aggregation and limits unique files, not chunks.
  • File ordering is deterministic, including deterministic tie-breaking.
  • Path filters are applied before file aggregation and continue to work in both modes.
  • Selected file results load complete textual file content through the existing source-loading boundary.
  • noContent omits result text in both modes and prevents content hydration.
  • contextLines has no effect in file mode.
  • maxCharacters remains a post-hydration output budget in both modes and can truncate only the last emitted result.
  • Truncated results retain the existing [...] marker and expose truncated: true; complete results do not expose the marker or flag.
  • Existing chunk-mode ranking and output behavior remains unchanged apart from the additive truncation signal.
  • Clipboard copy retains formatted text and works for file-mode results without copying JSON.
  • File aliases persist their mode and cannot override it at run time.
  • Alias schema validation rejects persisted entries that do not contain the required mode.
  • CLI, alias execution, and MCP tests cover the same mode semantics.
  • Tests cover single-hit files, multiple chunks from one file, duplicate elimination, top-after-aggregation, ties, path filters, full-file hydration, noContent, contextLines, maxCharacters, truncation, and clipboard output.

Testing Decisions

  • Test through public domain, application, command, alias, and MCP interfaces; do not assert private helper structure or implementation-specific grouping loops.
  • Extend query schema tests to cover the mode contract, mode-specific result shapes, optional content, and the truncation signal.
  • Extend application query tests with fixtures where several chunks belong to one file and where the strongest chunk is not the first chunk encountered by a naive top limit. Assert unique-file ranking and top-after-aggregation behavior.
  • Test path filtering before aggregation so ignored files cannot reappear through another chunk.
  • Test full-file hydration against the existing in-memory filesystem/test layers and verify that file mode returns source outside the matching chunk range.
  • Extend output-format tests for both modes, including no-content metadata-only output, ignored context lines, complete content, budget truncation, and the machine-readable truncated marker.
  • Preserve and extend existing command tests for formatted human output and clipboard copy; assert that copy remains formatted text rather than JSON.
  • Extend alias store and alias application tests to assert persisted mode, stable mode during alias execution, and rejection of entries without mode.
  • Extend MCP contract/handler tests to assert that the query and alias tools expose and preserve mode semantics.
  • Add adversarial retrieval fixtures for one strong chunk versus several mediocre chunks, duplicate/overlapping fallback chunks, equal scores, long files, and empty files. The MVP must remain bounded, deterministic, and free of unbounded long-file accumulation.

Out of Scope

  • Replacing chunk retrieval with file-level embedding or file-level indexing.
  • Changing Dense, Sparse, BM25, identifier, CamelCase, DBSF, evidence-router, or RRF behavior.
  • Unbounded score summation, average scoring, or a tuned multi-hit bonus in the MVP.
  • Returning matching chunks, line spans, context windows, or explanations inside file results.
  • Adding separate indexing or ranking for empty files.
  • Adding a hidden maximum file-size policy.
  • Adding binary extraction, new content processors, or changing which files are indexable.
  • Changing the existing clipboard output from formatted text to JSON.
  • Supporting a mode override when running a saved alias.
  • Backward-compatible alias migration or automatic healing of entries without mode.
  • Splitting this feature into ceremonial sub-issues; this issue is intentionally the complete, small feature slice.

Further Notes

The feature uses the existing distinction between chunk retrieval and lazy source loading: chunks remain the relevance evidence, while the file result is the requested output granularity. The MVP uses the strongest fused chunk score (max) for each file. Supporting-hit bonuses are out of scope until a later benchmark-backed decision.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestready-for-agentFully specified, ready for an AFK agent

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions