HyperMnesia is one layer of a small stack. It owns knowledge that persists — docs, the architectural map, and personal memory. It deliberately does not index code symbols; that job belongs to a live language-server layer it sits next to. This page describes the whole picture and how the pieces divide the work.
flowchart TB
AGENT["🤖 AI coding agent<br/>(Claude Code / any MCP client)"]
subgraph code["⚙️ Code layer · live, no index to maintain"]
SER["Serena (or any LSP-backed symbol MCP)"]
LSP["language servers<br/>pyright · gopls · rust-analyzer · tsserver · ..."]
SER --- LSP
end
subgraph know["📚 Knowledge layer · HyperMnesia"]
HM["MCP server"]
PG[("Postgres + pgvector<br/>docs · map · mem.*")]
HM --- PG
end
AGENT -->|"where is symbol X? who calls it?"| SER
AGENT -->|"what rules apply here? where's the doc? what do I know?"| HM
SER -->|"live symbols, defs, refs"| AGENT
HM -->|"constraints · doc chunks · memories"| AGENT
- Code layer (LSP symbols). A language server already builds and maintains a precise index of your code — definitions, references, call hierarchy, types — and updates it as you type. Exposed to the agent through an MCP wrapper (we use Serena), it answers structural questions about code with zero staleness. HyperMnesia does not duplicate this: re-embedding source on every edit would be wasteful and always a step behind the LSP.
- Knowledge layer (HyperMnesia). Everything that is not live code structure and that you
want to persist and retrieve:
- docs (markdown) → chunked, embedded (bge-m3), searched by hybrid RRF (Tier 2);
- architectural map → a hand-authored component/constraint graph, resolved deterministically from a file path (Tier 0/1) — "what rules apply to this file, before I edit it";
- personal memory → durable, bi-temporal facts/preferences/decisions distilled from sessions and injected back.
- Agent. The MCP client (Claude Code or any) talks to both servers. Neither knows about the other; they compose in the client.
| Question the agent has | Answered by | Why there |
|---|---|---|
"Where is foo defined? Who calls it? What's its type?" |
Serena / LSP | live index, always fresh, no re-embed |
"What rules / invariants apply to src/api/routes.py?" |
HyperMnesia Tier 0/1 | deterministic map, no model in the loop |
| "Where's the doc explaining the auth flow?" | HyperMnesia Tier 2 | hybrid RRF over doc chunks |
| "What did the owner decide about deploys, and why?" | HyperMnesia memory | bi-temporal, supersede-not-overwrite |
| "What are this repo's components at a glance?" | HyperMnesia project map | Tier 0 |
Rule of thumb: live code → the LSP layer; anything you want to remember or that lives in prose → HyperMnesia. Code that is stable and worth explaining (architecture, invariants) is captured once in the map/docs; code that changes constantly stays with the LSP.
Both are just MCP servers. In Claude Code's .mcp.json:
{
"mcpServers": {
"hypermnesia": {
"command": "/opt/hypermnesia/mcp-server/target/release/hypermnesia-mcp",
"env": {
"HM_REPO": "myrepo",
"DATABASE_URL": "postgresql://hm:pass@localhost:5432/hypermnesia",
"EMBED_BACKEND": "ollama",
"HM_SEARCH": "/opt/hypermnesia/ingest/search.py",
"HM_MEM_OPS": "/opt/hypermnesia/ingest/mem_ops.py",
"HM_RERANK": "/opt/hypermnesia/rerank/search_reranked.py"
}
},
"serena": {
"command": "serena",
"args": ["start-mcp-server", "--context", "claude-code", "--project", "/path/to/myrepo"]
}
}
}Any LSP-backed symbol MCP works in place of Serena; HyperMnesia doesn't depend on which one.
Ingest (offline, per doc change): ingest_repo.py enumerates the repo's tracked markdown (git ls-files; --walk for a tree kept out of git), chunks by
heading, and emits SQL; embed_chunks.py fills chunks.embedding via Ollama/TEI. The
architectural map is seeded once by hand (examples/seed_example.sql), refreshed when the
architecture changes.
Retrieval (per agent request): the MCP server resolves a path to its component and returns constraints deterministically (Tier 0/1), or embeds a query and fuses vector + full-text via RRF, optionally reranked (Tier 2). Fails open: embedder down → lexical-only; DB down → the tool surfaces the error, never blocks the agent.
Constraint injection (per edit, unprompted): a PreToolUse hook (hooks/arch_invariants.py,
matcher Edit|Write|MultiEdit) resolves the file about to be edited to its component and injects
the applicable must invariants as additionalContext — so Tier 1 reaches the agent
deterministically, without it having to think to call the get_constraints tool. This is the
architectural-memory counterpart to the personal-memory recall hook; both make delivery, not just
storage, the point. Fail-open: no DB / no match → the hook stays silent and the edit proceeds.
Keeping the map honest: the hand-authored map is only trustworthy while it tracks the tree.
ci/freshness.py makes decay mechanical — it flags component globs matching no file (a moved file
that silently unhooked its constraints — the worst case, since the map then lies more confidently
than search would), docs ingested behind HEAD, and — informational only — how many constraints
have no source document, since a full re-ingest nulls those links (the FK is ON DELETE SET NULL)
and many seed rows legitimately carry none. Run it in
CI or on a schedule against each target repo; it exits non-zero on orphaned globs, and refuses
before any of those checks if no component is mapped under the scope it was given — every check is
scoped by repo, so a mistyped or renamed scope would otherwise report zero of everything and
pass.
Personal memory (background): capture hooks enqueue session transcripts; a scheduled job
distills them to memories via a pluggable LLM; a daily consolidator merges near-duplicates behind
a confidence gate + review queue; a weekly reflect pass synthesizes each project's SHARED memories into
one knowledge page, stored as a memory (metadata.kind='page') that recall can surface as a
single overview. Recall/profile hooks inject relevant memories into the prompt, scoped to the
reader — see MEMORY.md.
Configuration (shared): the tunables the passes read — thresholds, model names, batch sizes —
come from one file, ~/.claude/hypermnesia.env, loaded on import by hooks/_mem_common.py and
ingest/_common.py. A variable already in the environment beats it, so a one-off run needs no
edit. The loader refuses the file entirely unless it is the owner's private file in a private
directory, and accepts only an exact list of names: the file must not be able to choose an
interpreter, a destination or a credential. See the table in INSTALL.md.
Operator's view (optional): console/ is a menu-bar tray and four command-line tools over the
same store — volumes, scheduled jobs, the settings above, and a first-run wizard. It reaches the
database through one setting, a command that receives SQL on stdin, so it works against a local
psql, a container, or a pod without knowing which. It is the surface where this system's own
failures are meant to become visible; DIAGNOSTICS.md lists every complaint it
can print.
See DESIGN.md for the reasoning behind the tiers and MEMORY.md for the personal-memory model.
Worth spelling out, because it is the single easiest way to break a store like this silently.
Every vector in the database — document chunks and memories alike — must come from the same embedding model. Vectors from two models do not live in the same space, so a query embedded by one cannot find rows embedded by the other. Nothing errors. Search simply gets worse, in a way that reads as "the corpus does not cover that" rather than as a bug, and it stays broken until someone re-embeds everything.
A concrete deployment, as an example of how to satisfy that while still using the right hardware for each job:
| Serving (per query) | Bulk (re-index, onboarding a repo) | |
|---|---|---|
| Model | BAAI/bge-m3 |
BAAI/bge-m3 |
| Runtime | text-embeddings-inference 1.7, CPU, in-cluster | Ollama, Apple Silicon |
| Shape | float32, 1024 dims, 8192-token input | same model, F16 weights |
| Why here | one query at a time, must always be up | ~10 chunks/s; a bulk run on the CPU replica takes 8-9 cores and starves live queries |
The split exists because the CPU replica that answers queries is easily saturated: one bulk client pushed it to 8-9 cores and ordinary queries stopped fitting their timeout. Moving bulk work to a local GPU/Neural-Engine runtime keeps serving responsive.
The split is only safe because both sides run the same model, and that is worth verifying rather than assuming. Take a few chunks that were embedded locally, re-embed their exact text through the serving runtime, and compare:
SELECT embedding <=> '[...vector from the serving runtime...]'::vector FROM chunks WHERE id = 52396;Cosine distance came back at ~1e-5 — float rounding, not a difference in meaning — so the vectors are interchangeable. A different model would have shown a distance in the tenths.
chunks.embedding_model and mem.memories.embedding_model record which model produced each
vector, so a silent swap becomes a visible mixture:
SELECT embedding_model, count(*) FROM chunks GROUP BY 1; -- expect exactly one rowIf that query ever returns two rows, retrieval quality is already degraded for part of the store, and the fix is to re-embed the minority — not to tune the ranking.
HyperMnesia is self-hosted memory + architectural control for coding agents, where you own the store. That's a different niche from managed consumer-chat memory. For context, Anthropic's Bringing memory to Claude (11 September 2025, updated 23 October 2025) gives Claude's own apps a memory that is optional rather than on by default, with a memory summary you can read and edit, memories scoped per project, incognito chats that are never remembered, and import/export. HyperMnesia differs on purpose: the store is your Postgres (inspectable and editable in SQL), it runs on local models with no data leaving your box, it is bi-temporal with supersede-not-overwrite + a review gate, and it is built around a coding agent's needs (doc-RAG + a deterministic constraint map), not chat. Complementary philosophies, different problems.
For the neighbours that solve this same problem, see COMPARISON.md.