Problem statement
Graph search already handles common conversational noise and several engineering
intents. What is still missing is one readable regression corpus proving that a
small set of realistic questions keeps returning the expected evidence near the
top as providers and ranking rules evolve.
Proposed solution
Current foundation
- deterministic local search with bounded output;
- natural-language stopword filtering and engineering-term normalization;
- rare-term, exact-label, alias, language, and architecture-intent weighting;
- project-subject qualifiers that stop generic architecture terms from
outranking an exact project or runtime subject;
- deterministic score tie-breaking and broad-query diversification;
- focused tests for authentication wording, generic service noise, long
architecture queries, language facets, and operational surfaces.
Remaining work
Create one compact table-driven fixture with at least five user questions. Keep
the expected entity IDs and maximum accepted rank next to each query so a ranking
change produces an understandable failure.
The initial questions should cover:
- authentication ownership;
- API/service ownership;
- dependency impact;
- CI/release workflow location;
- health endpoint location.
Acceptance criteria
Completion evidence
Publish the query table in the test failure output and document the focused local
command. If every criterion is already proven elsewhere on current main, close
#14 with links to those exact tests instead of adding duplicate fixtures.
Suggested starting points:
packages/cli/src/workspace-knowledge-graph-query.ts
packages/cli/src/__tests__/workspace-knowledge-graph.test.ts
packages/cli/scripts/real-world-qualification.mjs
Alternatives considered
The following alternatives or adjacent concerns are intentionally outside this issue:
- embeddings, remote semantic search, or model-based ranking;
- general provider semantic conformance, which is tracked separately;
- claiming relevance beyond the published fixture/query set.
Expected impact
Contributors get a small objective relevance target, while developers, IDEs,
MCP clients, and agents receive more repeatable bounded evidence.
Problem statement
Graph search already handles common conversational noise and several engineering
intents. What is still missing is one readable regression corpus proving that a
small set of realistic questions keeps returning the expected evidence near the
top as providers and ranking rules evolve.
Proposed solution
Current foundation
outranking an exact project or runtime subject;
architecture queries, language facets, and operational surfaces.
Remaining work
Create one compact table-driven fixture with at least five user questions. Keep
the expected entity IDs and maximum accepted rank next to each query so a ranking
change produces an understandable failure.
The initial questions should cover:
Acceptance criteria
Completion evidence
Publish the query table in the test failure output and document the focused local
command. If every criterion is already proven elsewhere on current
main, close#14 with links to those exact tests instead of adding duplicate fixtures.
Suggested starting points:
packages/cli/src/workspace-knowledge-graph-query.tspackages/cli/src/__tests__/workspace-knowledge-graph.test.tspackages/cli/scripts/real-world-qualification.mjsAlternatives considered
The following alternatives or adjacent concerns are intentionally outside this issue:
Expected impact
Contributors get a small objective relevance target, while developers, IDEs,
MCP clients, and agents receive more repeatable bounded evidence.