Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -49,3 +49,8 @@ jobs:
# Deterministic engine evals only — the agent suite costs tokens and is run manually.
- name: Run engine evals
run: npm run eval:engine

# The hard dataset's engine suite is equally deterministic, and it also guards the
# glossary -> term-resolution chain that the agent A/B depends on.
- name: Run engine evals (hard dataset)
run: npm run eval:engine:hard
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,30 @@ and public product updates.

## Unreleased

### A harder trap dataset, and a moat claim that now holds

- New `evals/dataset-hard/` (11 tables, deliberately bad naming) plus a committed glossary at
`evals/glossary-hard.json`. It exists because the first dataset stopped discriminating: 8 of
its 12 cases pass in *both* A/B arms
- Traps were chosen against the engine's **measured** solvable envelope, not intuition. An FK is
only discoverable when it shares a token with `"<keyTable> <keyColumn>"`, so `cust_ref` and
`owner` were rejected - the engine cannot solve those either, so they would fail in both arms
and discriminate nothing
- **The A/B now discriminates: grounded 29/36 runs (80.6%, 9/12 cases) vs raw-sql 21/36 (58.3%,
6/12), a +22.2 point delta with both metrics agreeing in direction**, versus +2.8 and
disagreeing on the original dataset. Mean tool steps 1.5 vs 4.5. Validity checks clean:
neither arm hit the turn budget, both baseline controls passed in both arms
- Every one of the control arm's failures has the same cause - it does not exclude void
invoices - which is exactly the kind of rule a schema cannot express and a glossary can
- Two results kept in the open rather than tuned away: grounding **lost** `revenue-by-category`
0/3 vs 1/3 by summing an invoice-grain measure at line grain, and both arms fail the fan-out
case identically (8,156 vs 2,039), so grounding does not prevent fan-out once the agent
hand-writes SQL instead of using `query_metric`
- `eval:engine:hard` is 25/25 and gates CI. Drop `--glossary` and exactly the five term cases
fail, which a test asserts - the enrichment chain cannot silently rot
- Every `expectedSql` was verified against the CSVs before being committed, and a test re-runs
all twelve so ground truth cannot drift

### Enrichment now reaches the agent and the resolver (bug fix)

- `enrich --apply` wrote descriptions and synonyms that **nothing ever read back**: there was no
Expand Down
67 changes: 52 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,11 +172,17 @@ The agent never sees a write path: `run_sql` is read-only-gated, and with `--db`
## `eval`: know whether a change made it better

```bash
npm run eval:engine # deterministic, no API key, runs in CI
npm run eval:agent # needs ANTHROPIC_API_KEY, costs tokens
npm run eval:ab # grounded vs raw-SQL, both arms interleaved
npm run eval:engine # deterministic, no API key, runs in CI
npm run eval:engine:hard # same, against the harder dataset
npm run eval:agent # needs ANTHROPIC_API_KEY, costs tokens
npm run eval:ab # grounded vs raw-SQL, both arms interleaved
npm run eval:ab:hard # the same A/B on the harder dataset
```

Any suite can be pointed anywhere: `--dataset <folder> --cases-file <path> --glossary <path>`.
Every report records which dataset, cases and curation produced it, so a number can never
be quoted without its setup.

Both suites score against `evals/dataset/` - a small e-commerce dataset built so a
careless answer is *wrong*, not just differently phrased:

Expand All @@ -188,9 +194,31 @@ careless answer is *wrong*, not just differently phrased:
| **Distinct vs count** | 9 customers placed 12 orders |
| **Null join** | one order has no line items; INNER vs LEFT changes the result |

`evals/dataset-hard/` is the second, harder dataset: 11 tables with deliberately bad naming,
plus a committed glossary (`evals/glossary-hard.json`) standing in for a user who has run
`enrich`. It exists because the first dataset stopped discriminating - 8 of its 12 cases now
pass in *both* A/B arms. Its traps were chosen against the engine's **measured** solvable
envelope rather than by intuition:

| Trap | Why grounding can win it |
|---|---|
| Opaque foreign keys | `acct.cust_id -> cust_master.id` is inferred at 92% from value overlap; a model reading names must guess which of eleven tables `cust_id` points at |
| Natural-key joins | lines reference products by `sku` and accounts reference regions by `region_cd`, never by `id` |
| Glossary term to measure | "revenue" reaches `sum_net_amt`; it shares no token with that name, so only the glossary connects them |
| Two money columns | `net_amt` is authoritative, `amt_txt` is a stale legacy mirror. Only the glossary says which |
| Soft delete and void | churned customers and void invoices must be excluded, a rule the schema cannot express |
| Decoy load buffer | `inv_staging` ids overlap `inv`; including it inflates revenue from 10,944 to 49,829 |
| Fan-out | tickets are a second `has_many` off customers, so a naive double join inflates one customer's revenue 4x |

Names like `cust_ref` or `owner` were **rejected** as traps: they score zero on name
similarity, so the engine cannot solve them either and they would fail in both arms while
discriminating nothing.

**Engine suite** asserts what the deterministic layers derive: inferred relationships and
their confidence, entity/dimension/measure derivation, metric compilation *including the
refusals that prevent fan-out*, and term resolution.
refusals that prevent fan-out*, and term resolution. It is 18/18 on the first dataset and
25/25 on the hard one; drop `--glossary` and exactly the five term cases fail, which is how
the enrichment chain is kept honest.

**Agent suite** puts each question through the real agent loop, then compares its rows
against ground truth produced by the case's `expectedSql` (never shown to the agent).
Expand All @@ -217,19 +245,28 @@ actually worth anything? It runs a `grounded` arm (what ships) against a `raw-sq
interleaved so API drift cannot masquerade as a result. The control is not crippled -
`SHOW`/`DESCRIBE` are read-only, so it discovers the schema itself.

The honest current answer, and the reason this suite exists:
The answer depends entirely on how hard the data is, which is the most useful thing this
suite has produced:

| | dataset | grounded | raw-sql | delta |
|---|---|---|---|---|
| run pass rate | original | 29/36 (80.6%) | 28/36 (77.8%) | +2.8 |
| run pass rate | **hard** | **29/36 (80.6%)** | **21/36 (58.3%)** | **+22.2** |
| mean tool steps | hard | **1.5** | 4.5 | ~60% fewer |

On the original 7-table dataset the grounding buys **nothing measurable on accuracy**: +2.8
points is inside the noise floor, the two metrics disagree on direction, and 8 of 12 cases pass
in both arms. A frontier model simply does not fall for a small fan-out trap.

| | grounded | raw-sql |
|---|---|---|
| run pass rate | 29/36 (80.6%) | 28/36 (77.8%) |
| strict cases | 8/12 | 9/12 |
| **mean tool steps** | **1.7** | **4.3** |
On the hard dataset it wins clearly, and for a legible reason: every one of the control arm's
failures traces to the same thing - it does not exclude void invoices, a business rule the
schema cannot express and only the glossary carries. Two cases go 3/3 versus **0/3**.

On **accuracy the grounding does not yet pay for itself** on this dataset: +2.8 points is
inside the noise floor, the two metrics disagree on direction, and 8 of 12 cases pass 3/3 in
*both* arms - a frontier model does not fall for a 7-table fan-out trap. On **efficiency it
clearly does**: ~60% fewer exploration steps, on every case. Both numbers are printed with
validity checks (turn-budget exhaustion, baseline controls) that must be read first.
Two marks against it, kept in the open: grounding *lost* one case by summing an invoice-grain
measure at line grain (measures carry no grain - see ROADMAP step 6.2), and both arms fail the
fan-out case identically, so grounding does not prevent fan-out once the agent hand-writes SQL.
Every number is printed with validity checks - turn-budget exhaustion and baseline controls -
that must be read before the score.

## `explain`: justify every join

Expand Down
102 changes: 89 additions & 13 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ the reason the understanding engine is built **CLI-first**.
**Surface decision (2026-07): terminal-first, desktop app as the flagship.**
The browser app was retired (tag `web-final`, ~6k LOC removed); repeating it in a TUI would compete with polished terminal SQL IDEs (Harlequin) on ground that is not our moat.
The surfaces today are the CLI and the MCP server.
The planned flagship surface (step 7 below, decided 2026-07-25) is a native macOS desktop app: a data cockpit that embeds a libghostty terminal running Claude Code / Codex under the user's own subscription (no BYOK), with native data panels driven by the MCP channel.
The CLI and MCP server are the app's foundation and stay first-class surfaces; the web only ever returns as an intro/landing page once the rename settles (step 8), never as a product UI.
The planned flagship surface (step 8 below, decided 2026-07-25) is a native macOS desktop app: a data cockpit that embeds a libghostty terminal running Claude Code / Codex under the user's own subscription (no BYOK), with native data panels driven by the MCP channel.
The CLI and MCP server are the app's foundation and stay first-class surfaces; the web only ever returns as an intro/landing page once the rename settles (step 9), never as a product UI.

> Cursor understands code → generates code → edits code → runs code.
> QueryPad understands datasets → infers relationships → generates SQL → executes analysis → explains findings.
Expand Down Expand Up @@ -175,13 +175,15 @@ architecture, naming). The intended moat is **semantic-first + local-first**: an
governed semantic model is reliable where naked text-to-SQL is not, and almost every funded
competitor is cloud/warehouse-native.

**Status of that claim, measured (2026-07-25, step 4.2):** the accuracy half is **not supported
yet**. Against a `run_sql`-only control on the trap dataset, grounding moved the run pass rate by
+2.8 points - well inside the noise floor - and the two arms disagreed on direction depending on the
metric. What grounding *did* buy, unambiguously, is **efficiency**: 1.7 versus 4.3 mean tool steps, a
~60% cut, on every single case. Treat "grounded is more accurate" as an open hypothesis needing a
harder dataset, and "grounded is cheaper and more auditable" as measured fact. Build one step at a
time:
**Status of that claim, measured twice (steps 4.2 and 4.3):** it holds, but only once the data is
hard enough to tell. On the original trap dataset grounding moved the run pass rate by +2.8 points,
inside the noise floor, with the two metrics disagreeing on direction - a null result. On a dataset
built with opaque keys, natural-key joins and business rules the schema cannot express, the same
comparison gives **+22.2 points (80.6% vs 58.3%)** with both metrics agreeing, and the control's
failures collapse to a single cause: it does not know the rules only a glossary carries.
Efficiency held in both runs (**1.5 vs 4.5** mean tool steps here, ~60% fewer). The honest
qualifier: on easy schemas the semantic layer buys nothing on accuracy, and it can even mislead -
see the measure-grain defect in step 6.2. Build one step at a time:

1. **Agentic `ask` loop** — self-correcting, tool-using. ✅ Built.
2. **Semantic layer** — ✅ Built (all five sub-steps). (Research-settled architecture: structured YAML core → DuckDB hybrid
Expand Down Expand Up @@ -248,14 +250,88 @@ time:
lever is the verification checklist, not more semantic layer. And the trap dataset needs to
get harder (more tables, worse names, genuinely ambiguous domains) before it can discriminate
accuracy at this model tier. Per `AGENTS.md` nothing was tuned to improve the number.
3. **Hard dataset** - ✅ Built (`evals/dataset-hard/`, 11 tables + a committed glossary at
`evals/glossary-hard.json`; `eval:engine:hard` gates CI, `eval:ab:hard` is manual).
Built to the *measured* solvable envelope rather than by intuition: an FK is only
discoverable when it shares a token with `"<keyTable> <keyColumn>"`
(`NAME_SIMILARITY_FLOOR`, `signals.ts:10`), so names like `cust_ref` or `owner` were
rejected as traps - the engine cannot solve them either, so they would fail in both arms
and discriminate nothing. The traps that remain sit in the band where the engine infers a
join at 92-100% from value overlap while a model reading `cust_id` against eleven tables
has to guess: opaque FKs, two natural-key joins (`sku`, `region_cd`), a soft-delete flag,
a void-status filter, a decoy load buffer whose ids overlap the real invoice table, a
stale legacy money column beside the authoritative one, and a second `has_many` that makes
fan-out possible.
**Two findings before a single agent case ran.** (a) The dataset immediately exposed a real
engine defect: `inv_line.unit_amt -> prod.list_amt` was inferred at 81% purely because
money values coincided with the 6 unique prices in a lookup table, and because a
relationship endpoint is excluded from measures, `inv.net_amt` silently stopped being a
measure at all. `isKeyCandidate` (`relationships.ts:38`) only asks whether the *target* is
unique and non-null, which any small price column satisfies - see step 6.1. The fixture was
made realistic (negotiated prices, not list prices) so the intended traps work; the defect
is recorded, not worked around. (b) The glossary chain is provably load-bearing: the hard
engine suite scores 25/25 with `--glossary` and exactly the five `hard-term-*` cases fail
without it, and a test asserts that.
**The A/B on this dataset discriminates.** Same configuration as before (12 cases,
repeat 3, verify on, maxSteps 12), validity checks clean (neither arm hit the turn budget;
both baseline controls 2/2 in both arms):

| | grounded | raw-sql |
|---|---|---|
| run pass rate | **29/36 (80.6%)** | 21/36 (58.3%) |
| strict cases | **9/12** | 6/12 |
| mean tool steps | **1.5** | 4.5 |

**+22.2 points**, and unlike the first dataset both metrics now agree in direction. The
raw-sql arm's failures share one root cause: it never excludes void invoices, so it
returns 11,275.50 / 1,297.50 / 2,535 / 1,128 where the answers are 10,944 / 1,020 /
2,257.50 / 1,074. That is a business rule the schema cannot express and only the glossary
carries, which is exactly what the semantic layer is for. Biggest gaps:
`hard-revenue-by-region` and `hard-top-customer` are 3/3 vs **0/3**.
**Two honest marks against it.** Grounding *lost* `hard-revenue-by-category` 0/3 vs 1/3
through the measure-grain defect above - the one case where being handed a measure was
worse than having none. And `hard-fanout-revenue-and-cases` is **0/3 in both arms**: both
agents write the naive double join and inflate one customer's revenue 4x (8,156 vs 2,039),
so grounding does not prevent fan-out once the agent leaves `query_metric` and hand-writes
SQL. Per `AGENTS.md` nothing was tuned after seeing these numbers.
5. **MCP server** — ✅ Built. `querypad mcp` serves the read-only toolkit over stdio. The
tools are not a reimplementation: `createDataToolkit` (`src/core/agent/toolkit.ts`) is
the single definition that both the internal `ask` loop and the MCP server consume, so
an external agent sees exactly the tools our own agent uses — plus `describe_dataset`,
which hands over the grounding context `ask` would otherwise put in its system prompt.
With the web app retired this is the interactive surface: the coding agent is the UI.
6. Short planning/decomposition for multi-part questions (bounded).
7. **Native desktop app** (decided 2026-07-25) — the flagship product surface.
6. **Engine defects surfaced by the hard dataset** - open, and worth fixing before more
semantic-layer work.
1. **Numeric value overlap creates phantom foreign keys.** `isKeyCandidate`
(`src/core/discovery/relationships.ts:38`) accepts any column that is unique and non-null
in its own table, so a small lookup table's `list_amt` is a valid FK *target*. Every money
column whose values coincide with those prices then gets an edge (measured: 81%, 68%, 54%
on the hard fixture before it was made realistic). The damage is not just a wrong edge in
the graph: `keyColumns` (`semantic-model.ts:75`) excludes both endpoints of every edge from
dimensions **and** measures, so the table's real money measure disappears without a word.
Candidate fix: require an FK target to look like a key (id-like name, or referenced by a
name-similar column), or refuse targets that are themselves measures.
2. **Measures have no grain, so naming one can actively mislead.** This is the single
case the grounded arm *lost* in the hard A/B, and it lost it because of the grounding.
The glossary names `inv.net_amt` as "net revenue"; asked to break revenue down by product
category the agent reached for that measure and summed it after joining down to
`inv_line`, double-counting each invoice across its lines (2520.5 instead of 1074 for
Software). The data is not ambiguous - line-level and header totals are both exactly
10,944 - so this is a real grain error, not a bad case. `SemanticMeasure` records
`agg` and `column` but nothing about the grain it is valid at, `query_metric` refuses
cross-grain joins only inside its own compiler, and `compile-metric.ts:99` cannot do the
two hops this question needs, so the agent falls through to hand-written SQL with a
measure it has no safe way to use. Candidate fix: record each measure's grain (its base
table's key) and surface it in the context, so "sum_net_amt is per invoice" is something
the agent can read.
3. **Duplicate measure names resolve silently.** Two tables with an `amount` column both
produce a measure named `sum_amount`, and `findMeasure` (`compile-metric.ts:33`) returns the
first by entity order. Same for duplicate dimension names, and `ensureJoin` matches on the
table pair rather than the column, so two FKs into one target pick whichever edge sorts
first. Deliberately left out of the hard dataset: it is an engine ambiguity to fix, not a
grounding trap to grade.
7. Short planning/decomposition for multi-part questions (bounded).
8. **Native desktop app** (decided 2026-07-25) — the flagship product surface.
A native macOS app (Swift + AppKit) embeds a libghostty terminal pane running
Claude Code / Codex under the user's own subscription (no BYOK).
The app owns the querypad engine as a bundled subprocess and exposes the MCP
Expand Down Expand Up @@ -284,8 +360,8 @@ time:
graph/chart panels, agent picker (claude / codex).
Constraints: macOS-only initially (the proven libghostty embedding path);
the Node engine ships bundled inside the .app (the native DuckDB addon rules
out easy single-binary compiles); distribution is gated on the rename (step 8).
8. **Rename** (package / bin / domain / README) — *gated on formal trademark + domain
out easy single-binary compiles); distribution is gated on the rename (step 9).
9. **Rename** (package / bin / domain / README) — *gated on formal trademark + domain
clearance* (the name "datapad" was rejected: it collides with an active, funded
competitor in the same category; "grain" has an npm squatter + a language collision).
The artifact dir is already brand-independent (`.datactx/`), and npm publish waits
Expand Down
Loading