Skip to content

Latest commit

 

History

History
201 lines (163 loc) · 10.8 KB

File metadata and controls

201 lines (163 loc) · 10.8 KB

System architecture

HyperMnesia is one layer of a small stack. It owns knowledge that persists — docs, the architectural map, and personal memory. It deliberately does not index code symbols; that job belongs to a live language-server layer it sits next to. This page describes the whole picture and how the pieces divide the work.

The three layers

flowchart TB
    AGENT["🤖 AI coding agent<br/>(Claude Code / any MCP client)"]

    subgraph code["⚙️ Code layer · live, no index to maintain"]
        SER["Serena (or any LSP-backed symbol MCP)"]
        LSP["language servers<br/>pyright · gopls · rust-analyzer · tsserver · ..."]
        SER --- LSP
    end

    subgraph know["📚 Knowledge layer · HyperMnesia"]
        HM["MCP server"]
        PG[("Postgres + pgvector<br/>docs · map · mem.*")]
        HM --- PG
    end

    AGENT -->|"where is symbol X? who calls it?"| SER
    AGENT -->|"what rules apply here? where's the doc? what do I know?"| HM
    SER -->|"live symbols, defs, refs"| AGENT
    HM -->|"constraints · doc chunks · memories"| AGENT
Loading
  • Code layer (LSP symbols). A language server already builds and maintains a precise index of your code — definitions, references, call hierarchy, types — and updates it as you type. Exposed to the agent through an MCP wrapper (we use Serena), it answers structural questions about code with zero staleness. HyperMnesia does not duplicate this: re-embedding source on every edit would be wasteful and always a step behind the LSP.
  • Knowledge layer (HyperMnesia). Everything that is not live code structure and that you want to persist and retrieve:
    • docs (markdown) → chunked, embedded (bge-m3), searched by hybrid RRF (Tier 2);
    • architectural map → a hand-authored component/constraint graph, resolved deterministically from a file path (Tier 0/1) — "what rules apply to this file, before I edit it";
    • personal memory → durable, bi-temporal facts/preferences/decisions distilled from sessions and injected back.
  • Agent. The MCP client (Claude Code or any) talks to both servers. Neither knows about the other; they compose in the client.

Division of labor

Question the agent has Answered by Why there
"Where is foo defined? Who calls it? What's its type?" Serena / LSP live index, always fresh, no re-embed
"What rules / invariants apply to src/api/routes.py?" HyperMnesia Tier 0/1 deterministic map, no model in the loop
"Where's the doc explaining the auth flow?" HyperMnesia Tier 2 hybrid RRF over doc chunks
"What did the owner decide about deploys, and why?" HyperMnesia memory bi-temporal, supersede-not-overwrite
"What are this repo's components at a glance?" HyperMnesia project map Tier 0

Rule of thumb: live code → the LSP layer; anything you want to remember or that lives in prose → HyperMnesia. Code that is stable and worth explaining (architecture, invariants) is captured once in the map/docs; code that changes constantly stays with the LSP.

Pairing them (one MCP client, two servers)

Both are just MCP servers. In Claude Code's .mcp.json:

{
  "mcpServers": {
    "hypermnesia": {
      "command": "/opt/hypermnesia/mcp-server/target/release/hypermnesia-mcp",
      "env": {
        "HM_REPO": "myrepo",
        "DATABASE_URL": "postgresql://hm:pass@localhost:5432/hypermnesia",
        "EMBED_BACKEND": "ollama",
        "HM_SEARCH":  "/opt/hypermnesia/ingest/search.py",
        "HM_MEM_OPS": "/opt/hypermnesia/ingest/mem_ops.py",
        "HM_RERANK":  "/opt/hypermnesia/rerank/search_reranked.py"
      }
    },
    "serena": {
      "command": "serena",
      "args": ["start-mcp-server", "--context", "claude-code", "--project", "/path/to/myrepo"]
    }
  }
}

Any LSP-backed symbol MCP works in place of Serena; HyperMnesia doesn't depend on which one.

Data flow

Ingest (offline, per doc change): ingest_repo.py enumerates the repo's tracked markdown (git ls-files; --walk for a tree kept out of git), chunks by heading, and emits SQL; embed_chunks.py fills chunks.embedding via Ollama/TEI. The architectural map is seeded once by hand (examples/seed_example.sql), refreshed when the architecture changes.

Retrieval (per agent request): the MCP server resolves a path to its component and returns constraints deterministically (Tier 0/1), or embeds a query and fuses vector + full-text via RRF, optionally reranked (Tier 2). Fails open: embedder down → lexical-only; DB down → the tool surfaces the error, never blocks the agent.

Constraint injection (per edit, unprompted): a PreToolUse hook (hooks/arch_invariants.py, matcher Edit|Write|MultiEdit) resolves the file about to be edited to its component and injects the applicable must invariants as additionalContext — so Tier 1 reaches the agent deterministically, without it having to think to call the get_constraints tool. This is the architectural-memory counterpart to the personal-memory recall hook; both make delivery, not just storage, the point. Fail-open: no DB / no match → the hook stays silent and the edit proceeds.

Keeping the map honest: the hand-authored map is only trustworthy while it tracks the tree. ci/freshness.py makes decay mechanical — it flags component globs matching no file (a moved file that silently unhooked its constraints — the worst case, since the map then lies more confidently than search would), docs ingested behind HEAD, and — informational only — how many constraints have no source document, since a full re-ingest nulls those links (the FK is ON DELETE SET NULL) and many seed rows legitimately carry none. Run it in CI or on a schedule against each target repo; it exits non-zero on orphaned globs, and refuses before any of those checks if no component is mapped under the scope it was given — every check is scoped by repo, so a mistyped or renamed scope would otherwise report zero of everything and pass.

Personal memory (background): capture hooks enqueue session transcripts; a scheduled job distills them to memories via a pluggable LLM; a daily consolidator merges near-duplicates behind a confidence gate + review queue; a weekly reflect pass synthesizes each project's SHARED memories into one knowledge page, stored as a memory (metadata.kind='page') that recall can surface as a single overview. Recall/profile hooks inject relevant memories into the prompt, scoped to the reader — see MEMORY.md.

Configuration (shared): the tunables the passes read — thresholds, model names, batch sizes — come from one file, ~/.claude/hypermnesia.env, loaded on import by hooks/_mem_common.py and ingest/_common.py. A variable already in the environment beats it, so a one-off run needs no edit. The loader refuses the file entirely unless it is the owner's private file in a private directory, and accepts only an exact list of names: the file must not be able to choose an interpreter, a destination or a credential. See the table in INSTALL.md.

Operator's view (optional): console/ is a menu-bar tray and four command-line tools over the same store — volumes, scheduled jobs, the settings above, and a first-run wizard. It reaches the database through one setting, a command that receives SQL on stdin, so it works against a local psql, a container, or a pod without knowing which. It is the surface where this system's own failures are meant to become visible; DIAGNOSTICS.md lists every complaint it can print.

See DESIGN.md for the reasoning behind the tiers and MEMORY.md for the personal-memory model.

One model, two runtimes: the embedder

Worth spelling out, because it is the single easiest way to break a store like this silently.

Every vector in the database — document chunks and memories alike — must come from the same embedding model. Vectors from two models do not live in the same space, so a query embedded by one cannot find rows embedded by the other. Nothing errors. Search simply gets worse, in a way that reads as "the corpus does not cover that" rather than as a bug, and it stays broken until someone re-embeds everything.

A concrete deployment, as an example of how to satisfy that while still using the right hardware for each job:

Serving (per query) Bulk (re-index, onboarding a repo)
Model BAAI/bge-m3 BAAI/bge-m3
Runtime text-embeddings-inference 1.7, CPU, in-cluster Ollama, Apple Silicon
Shape float32, 1024 dims, 8192-token input same model, F16 weights
Why here one query at a time, must always be up ~10 chunks/s; a bulk run on the CPU replica takes 8-9 cores and starves live queries

The split exists because the CPU replica that answers queries is easily saturated: one bulk client pushed it to 8-9 cores and ordinary queries stopped fitting their timeout. Moving bulk work to a local GPU/Neural-Engine runtime keeps serving responsive.

The split is only safe because both sides run the same model, and that is worth verifying rather than assuming. Take a few chunks that were embedded locally, re-embed their exact text through the serving runtime, and compare:

SELECT embedding <=> '[...vector from the serving runtime...]'::vector FROM chunks WHERE id = 52396;

Cosine distance came back at ~1e-5 — float rounding, not a difference in meaning — so the vectors are interchangeable. A different model would have shown a distance in the tenths.

chunks.embedding_model and mem.memories.embedding_model record which model produced each vector, so a silent swap becomes a visible mixture:

SELECT embedding_model, count(*) FROM chunks GROUP BY 1;   -- expect exactly one row

If that query ever returns two rows, retrieval quality is already degraded for part of the store, and the fix is to re-embed the minority — not to tune the ranking.

Related work

HyperMnesia is self-hosted memory + architectural control for coding agents, where you own the store. That's a different niche from managed consumer-chat memory. For context, Anthropic's Bringing memory to Claude (11 September 2025, updated 23 October 2025) gives Claude's own apps a memory that is optional rather than on by default, with a memory summary you can read and edit, memories scoped per project, incognito chats that are never remembered, and import/export. HyperMnesia differs on purpose: the store is your Postgres (inspectable and editable in SQL), it runs on local models with no data leaving your box, it is bi-temporal with supersede-not-overwrite + a review gate, and it is built around a coding agent's needs (doc-RAG + a deterministic constraint map), not chat. Complementary philosophies, different problems.

For the neighbours that solve this same problem, see COMPARISON.md.