Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions benchmark/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,4 @@ long time or incur inference cost.
- [`locomo/`](locomo/README.md): conversation-memory retrieval and end-to-end question-answer accuracy.
- [`locomo_plus/`](locomo_plus/README.md): pinned LoCoMo-Plus factual and cognitive memory evaluation, with a
four-case smoke profile, a bundled ten-case dataset with complete histories, and an explicit full profile.
- [`memory_capacity/`](memory_capacity/README.md): manifest capacity, append cost, and tombstone compaction without model calls.
70 changes: 70 additions & 0 deletions benchmark/memory_capacity/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Memory capacity benchmark

This provider-free benchmark appends one entry per Revision at 200, 1,000 and 5,000 entries, then retires 80% of the
entries, advances the tombstone recovery window, previews and commits compaction, and appends once more. It measures
canonical bytes, database bytes, mean append latency, the final 100 appends, affected projection rows, and preservation
of FTS hit identities across compaction. It retains historical manifests and entry bodies.

```bash
uv run python -m benchmark.memory_capacity --output .artifacts/memory-capacity/run-01/sqlite.json
uv run python -m benchmark.memory_capacity --backend oceanbase --output .artifacts/memory-capacity/run-01/oceanbase.json
```

Choose a new run directory for each measurement. Raw JSON, logs, and databases are local acceptance artifacts under the
Git-ignored `.artifacts/` directory. Attach reviewed results to the relevant PR or CI run and keep reusable methodology
and concise measurement summaries in this README.

SQLite creates a fresh database beside the output JSON, retained for inspection. Samples checkpoint and truncate the
WAL; the compaction report also records file bytes after `VACUUM`.
The full run requires several GB of disk because every historical manifest remains authoritative.

OceanBase requires `POWERCONTEXT_TEST_OCEANBASE_URL` pointing at a disposable test database. It creates a unique Scope;
reported database bytes cover the entire database, so use an otherwise idle database. The default observation delay is
30 seconds, adjustable with `--reclamation-delay`. No storage-engine compaction is forced. An unavailable database is
recorded as `not_run`, never represented as a measured result. No model calls are made on either backend.

OceanBase database bytes sum `OCCUPY_SIZE` from `oceanbase.DBA_OB_TABLE_SPACE_USAGE` for the selected database, including
its index tables. This measures reported SSTable occupancy, excluding memtables, transaction logs, and preallocated
cluster files; early samples can be zero before a flush. `information_schema.tables` statistics can remain zero even
after SSTables occupy space and are not used for this measurement.

`reclaimed_bytes` compares full canonical contents, including the new compaction audit records. The follow-up Revision
shows the continuing cost of the reduced manifest. Projection rows are counted using the database cursor's affected
row count; zero-row deletes do not count. FTS and active-head rows are included; immutable entry bodies are separate.

The latency sample is a local calibration, not an SLA. Default budgets remain 5,000 active entries, 10,000 manifest
entries, and 4 MiB pending deployment-specific latency requirements.


## Recorded local SQLite run

Windows, Python 3.11.9, one append per Revision; other repository validation ran concurrently, so these latencies are
indicative and must not be used as an isolated performance baseline. Byte counts are exact.

| Entries | Canonical bytes | Checkpointed database bytes | Mean append (ms) | Final 100 mean (ms) | Projection rows/append |
| --- | --- | --- | --- | --- | --- |
| 200 | 44,866 | 5,607,424 | 14.42 | 15.85 | 2 |
| 1,000 | 223,266 | 114,847,744 | 29.29 | 46.09 | 2 |
| 5,000 | 1,115,266 | 2,804,195,328 | 116.19 | 217.74 | 2 |

Compaction removed 4,000 tombstones with zero projection writes and preserved the same sentinel search hit. Complete
canonical content fell from 1,123,307 to 939,091 bytes in the compaction Revision, then to 223,489 bytes on the next
append. The audit delta accounts for the difference. Database bytes after that append were 2,815,942,656 and after
`VACUUM` were 2,815,492,096: old manifests still occupy live pages.

## OceanBase validation status

On 2026-09-26, all 13 OceanBase capacity behavior cases passed against OceanBase CE 4.3.5.6 without skips. Another
67 SQLite, HTTP/API contract, and projection regression checks passed. The complete 5,000-entry scale run and
compaction cycle also passed locally: 4,000 tombstones were removed, zero projection rows were written by compaction,
and the sentinel FTS hit identity was preserved.

| Entries | Canonical bytes | Observed database bytes | Mean append (ms) | Final 100 mean (ms) | Projection rows/append |
| --- | ---: | ---: | ---: | ---: | ---: |
| 200 | 44,866 | 73,706 | 24.69 | 25.94 | 1 |
| 1,000 | 223,266 | 73,706 | 36.14 | 47.70 | 1 |
| 5,000 | 1,115,266 | 1,679,697,682 | 99.85 | 180.02 | 1 |

The compaction revision reduced manifest entries from 5,000 to 1,000 and canonical bytes from 1,123,307 to 939,091;
the follow-up append produced 1,001 entries and 223,489 bytes. Database bytes stayed at 1,770,500,275 during the
compaction observation window, as historical revisions remain retained.
15 changes: 15 additions & 0 deletions benchmark/memory_capacity/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Copyright (c) 2026 OceanBase.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

"""Capacity envelope benchmark without model calls."""
211 changes: 211 additions & 0 deletions benchmark/memory_capacity/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,211 @@
# Copyright (c) 2026 OceanBase.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

from __future__ import annotations

import argparse
import asyncio
import json
import os
import platform
from pathlib import Path
from statistics import mean
from time import perf_counter
from uuid import uuid4

from pydantic import SecretStr
from sqlalchemy import event, text

from powercontext.builtin.artifacts.memory import MemoryEntryInput
from powercontext.builtin.artifacts.memory.canonical import memory_content_bytes
from powercontext.builtin.persistence.oceanbase import OceanBaseConfig
from powercontext.builtin.persistence.sqlite import SQLiteConfig
from powercontext.builtin.runtime import BuiltinConfig, open_builtin_contexts
from powercontext.builtin.runtime.config import RuntimeConfig


async def run(args): # noqa: C901
if args.backend == "oceanbase":
url = os.environ.get("POWERCONTEXT_TEST_OCEANBASE_URL")
if not url:
return {
"backend": "oceanbase",
"status": "not_run",
"reason": "POWERCONTEXT_TEST_OCEANBASE_URL is not configured",
}
database = OceanBaseConfig(url=SecretStr(url))
database_path = None
else:
database_path = args.output.parent / ("memory-capacity-" + uuid4().hex + ".db")
database = SQLiteConfig(url=f"sqlite+aiosqlite:///{database_path.resolve().as_posix()}")
config = BuiltinConfig(database=database, runtime=RuntimeConfig(memory_compaction_enabled=True))
output = {
"backend": args.backend,
"status": "running",
"python": platform.python_version(),
"platform": platform.platform(),
"counts": args.counts,
"final_window": args.final_window,
"measurements": [],
"reclamation": "WAL checkpoint before samples; VACUUM measured separately"
if database_path
else f"observed table bytes after {args.reclamation_delay}s; no forced engine compaction",
}
async with open_builtin_contexts(config) as contexts:
scope_id = "capacity-benchmark-" + uuid4().hex
service = (await contexts.get(scope_id)).artifacts.memory
projection_rows = 0

def record(_connection, cursor, statement, _parameters, _context, _executemany):
nonlocal projection_rows
sql = statement.lower().lstrip()
if sql.startswith(("insert", "update", "delete")) and any(
name in sql for name in ("pc_memory_entry_heads", "pc_memory_entry_fts")
):
projection_rows += max(cursor.rowcount, 0)

event.listen(contexts.database.engine.sync_engine, "after_cursor_execute", record)

async def database_bytes():
if database_path is not None:
async with contexts.database.engine.connect() as connection:
await connection.exec_driver_sql("PRAGMA wal_checkpoint(TRUNCATE)")
return database_path.stat().st_size
async with contexts.database.connection() as connection:
# General table statistics can remain zero while OceanBase already occupies SSTable space.
return int(
await connection.scalar(
text(
"SELECT COALESCE(SUM(OCCUPY_SIZE), 0) "
"FROM oceanbase.DBA_OB_TABLE_SPACE_USAGE WHERE DATABASE_NAME = DATABASE()"
)
)
or 0
)

async def snapshot(memory):
capacity = await service.capacity(memory)
return {
"revision": memory.revision,
"active_entries": capacity.active_entry_count,
"manifest_entries": capacity.manifest_entry_count,
"manifest_bytes": capacity.manifest_bytes,
"database_bytes": await database_bytes(),
}

memory = None
latencies = []
for number in range(1, max(args.counts) + 1):
started = perf_counter()
memory = await service.remember(
memory=memory,
entries=(
MemoryEntryInput(
kind="fact",
text=f"Capacity benchmark record {number}; project token item{number}.",
),
),
mode="append",
)
latencies.append((perf_counter() - started) * 1000)
if number in args.counts:
sample = await snapshot(memory)
sample.update({
"mean_append_ms": mean(latencies),
"mean_final_window_append_ms": mean(latencies[-args.final_window :]),
"projection_rows_per_append": projection_rows / number,
})
output["measurements"].append(sample)
print(json.dumps(sample), flush=True)
args.output.write_text(json.dumps(output, indent=2) + "\n", encoding="utf-8")
assert memory is not None # noqa: S101
entries = await service.entries(memory)
retired_entries = entries[: int(len(entries) * 0.8)]
retained = entries[-1]
memory = await service.forget(memory, entries=retired_entries)
for number in range(config.runtime.memory_compaction_min_tombstone_revisions):
memory = await service.remember(
memory=memory,
entries=(
MemoryEntryInput(
kind="fact",
text=f"Capacity sentinel retained record generation {number}.",
entry=retained,
),
),
mode="append",
)
assert memory is not None # noqa: S101
retained = next(entry for entry in await service.entries(memory) if entry.entry_id == retained.entry_id)
assert memory is not None # noqa: S101
before = await snapshot(memory)
hits_before = await service.search("Capacity sentinel retained", memories=(memory,), mode="fts")
preview = await service.compact(memory, dry_run=True)
writes_before = projection_rows
compacted = await service.compact(memory)
compaction_writes = projection_rows - writes_before
hits_after = await service.search("Capacity sentinel retained", memories=(compacted.memory,), mode="fts")
identities_before = [(hit.entry_id, hit.entry_version_id) for hit in hits_before.hits]
identities_after = [(hit.entry_id, hit.entry_version_id) for hit in hits_after.hits]
output["compaction"] = {
"before": before,
"after": await snapshot(compacted.memory),
"removed_entries": len(compacted.entry_ids),
"reclaimed_bytes": compacted.reclaimed_bytes,
"preview_matches": preview.entry_ids == compacted.entry_ids,
"projection_rows_written": compaction_writes,
"search_hit_count_before": len(identities_before),
"search_hit_count_after": len(identities_after),
"search_identities_preserved": identities_before == identities_after,
}
if not identities_before or identities_before != identities_after or compaction_writes:
raise RuntimeError("capacity benchmark invariants failed") # noqa: TRY003
followup = await service.remember(
memory=compacted.memory,
entries=(MemoryEntryInput(kind="fact", text="New entry after compaction."),),
mode="append",
)
assert followup is not None # noqa: S101
output["compaction"]["followup"] = await snapshot(followup)
if database_path is not None:
async with contexts.database.engine.connect() as connection:
await connection.exec_driver_sql("VACUUM")
output["compaction"]["post_vacuum_database_bytes"] = await database_bytes()
else:
await asyncio.sleep(args.reclamation_delay)
output["compaction"]["delayed_database_bytes"] = await database_bytes()
output["compaction"]["followup_canonical_bytes"] = len(memory_content_bytes(followup.content))
event.remove(contexts.database.engine.sync_engine, "after_cursor_execute", record)
output["status"] = "completed"
return output


def main():
parser = argparse.ArgumentParser(description="Measure the Memory capacity envelope without inference.")
parser.add_argument("--backend", choices=("sqlite", "oceanbase"), default="sqlite")
parser.add_argument("--counts", nargs="+", type=int, default=[200, 1000, 5000])
parser.add_argument("--final-window", type=int, default=100)
parser.add_argument("--reclamation-delay", type=float, default=30)
parser.add_argument("--output", type=Path, required=True)
args = parser.parse_args()
if min(args.counts) < 5 or max(args.counts) > 5000 or args.final_window < 1 or args.reclamation_delay < 0:
parser.error("counts must be 5..5000, final-window positive, and reclamation-delay nonnegative")
args.output.parent.mkdir(parents=True, exist_ok=True)
result = asyncio.run(run(args))
args.output.write_text(json.dumps(result, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")
print(args.output)


if __name__ == "__main__":
main()
Loading
Loading