Skip to content

[Feature]: Add a table-driven natural-language graph relevance corpus #14

Description

@Baziar

Problem statement

Graph search already handles common conversational noise and several engineering
intents. What is still missing is one readable regression corpus proving that a
small set of realistic questions keeps returning the expected evidence near the
top as providers and ranking rules evolve.

Proposed solution

Current foundation

  • deterministic local search with bounded output;
  • natural-language stopword filtering and engineering-term normalization;
  • rare-term, exact-label, alias, language, and architecture-intent weighting;
  • project-subject qualifiers that stop generic architecture terms from
    outranking an exact project or runtime subject;
  • deterministic score tie-breaking and broad-query diversification;
  • focused tests for authentication wording, generic service noise, long
    architecture queries, language facets, and operational surfaces.

Remaining work

Create one compact table-driven fixture with at least five user questions. Keep
the expected entity IDs and maximum accepted rank next to each query so a ranking
change produces an understandable failure.

The initial questions should cover:

  • authentication ownership;
  • API/service ownership;
  • dependency impact;
  • CI/release workflow location;
  • health endpoint location.

Acceptance criteria

  • At least five realistic queries and expected entity IDs share one fixture.
  • Every primary expected entity appears within the top three results.
  • Conversational filler cannot outrank the specific engineering terms.
  • Duplicate logical labels/aliases cannot consume the bounded result budget.
  • Two identical runs return the same entity order and scores.
  • Returned expected entities retain proof references where the fixture provides proof.
  • The suite is deterministic, offline, and passes in Linux/macOS/Windows CI.
  • Public graph/search contracts remain backward compatible.

Completion evidence

Publish the query table in the test failure output and document the focused local
command. If every criterion is already proven elsewhere on current main, close
#14 with links to those exact tests instead of adding duplicate fixtures.

Suggested starting points:

  • packages/cli/src/workspace-knowledge-graph-query.ts
  • packages/cli/src/__tests__/workspace-knowledge-graph.test.ts
  • packages/cli/scripts/real-world-qualification.mjs

Alternatives considered

The following alternatives or adjacent concerns are intentionally outside this issue:

  • embeddings, remote semantic search, or model-based ranking;
  • general provider semantic conformance, which is tracked separately;
  • claiming relevance beyond the published fixture/query set.

Expected impact

Contributors get a small objective relevance target, while developers, IDEs,
MCP clients, and agents receive more repeatable bounded evidence.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

area: graphGraph providers, identity, relations, proof, and querygood first issueGood for newcomershelp wantedExtra attention is needed

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions