The memory-tier counterpart to a doc-RAG recall@k score. Instead of a number, it asserts the
properties a bi-temporal memory store must have or it is silently wrong — the LongMemEval failure
modes, checked mechanically against the live serving path (whatever DATABASE_URL +
EMBED_BACKEND point at). Self-cleaning: probe rows are tagged metadata.probe=true and
hard-deleted afterwards.
DATABASE_URL=... EMBED_BACKEND=ollama python3 eval/mem_probes.py # exit 1 if any probe failsProbes:
- staleness — a superseded fact must not surface in default search; the new one must; the old
one stays reachable with
include_inactive. - abstention — an unknown-topic query must return nothing, not top-k noise.
- temporal — a fact whose
valid_tois in the past is hidden by default, visible withinclude_inactive. - recall — a paraphrased query finds a just-written distinctive fact in the top-3.
Run it after any change to the memory retrieval path. It needs an embedder (Ollama/TEI) and psql
on PATH; point HM_PYTHON / HM_MEM_OPS at your interpreter / mem_ops.py if not the defaults.
Running this against a populated store surfaced two real weaknesses, both now fixed in
ingest/mem_ops.py and worth recording so the design intent survives:
- The lexical leg had no relevance floor. The hybrid query is OR-converted (deliberately, so
identifiers and unstemmed forms match), and the composite
tsvectorincludes asimple(no-stopword) component. So a query sharing one incidental token with a memory — a stopword likeна/the, or a content word likehigh— produced a lexical hit that bypassed the semantic distance gate entirely. A Russian "recipe for borscht…" returned five unrelated infra memories. Fix: a lexical-only hit is kept only if it is also semantically plausible (embedding <=> query < MEM_LEX_MAXDIST, which defaults toMEM_SEM_MAXDISTitself — no slack, since unrelated memories start right above the gate). The lexical leg's job is to rescue near-misses of the semantic gate, not to admit strangers. The simple-leg query also drops stopwords and ≤2-char tokens to cut ranking noise. - The semantic gate was loose for bge-m3. Measured on a populated store: genuine paraphrase
recall tops out ~0.45 (8 queries: 0.34–0.45), unrelated queries start ~0.52 (6 queries:
0.52–0.68). The old default of 0.6 let topically-adjacent-but-unrelated queries through as
noise.
MEM_SEM_MAXDISTnow defaults to 0.5, in the measured gap; tune it per embedder with the same kind of measurement (mem_ops nearestgives the distance).
The abstention probe uses a wholly orthogonal topic, so it proves only that an unrelated query abstains. It deliberately does NOT exercise the near-topic case the distance tuning above addresses: a query that shares one incidental token with a memory is not covered by any probe today.
Latency, per stage, against your own store: the PreToolUse hook end to end (the one that runs on
every edit, process start included, because that is what you actually wait for), one query
embedding, the fused RRF statement, search.py end to end, and the reranker. Median, p90 and max
of N runs with the first discarded — a first run measures cold page cache, not this system.
DATABASE_URL=... ./hm latency # or: -n 20 -q "your own query"The corpus it ran against is printed once, as a header above the table — because a latency figure without a document and chunk count cannot be read at all, let alone compared.
Whether any of this makes an agent write better code. Doing that honestly needs a fixed task set, runs with and without injection, and control of everything else that differs between them. Nothing here does that, so no number about it appears anywhere in this repo.
The one retrieval-quality figure that does appear — the reranker moving recall@1 by +0.16 — comes
from a small private evaluation set, and is quoted as what it is: a result on one corpus, not a
benchmark. mem_probes.py above is asserted properties rather than a score for the same reason:
properties survive being run on a different corpus, and a score does not.