You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Compression vs Raw Data (LongMemEval, 50-entry sample)
Reference
Size
→ v3.12 AMS
Ratio
UTF-8 original text
1.21 MB
2.23 MB
0.54× (expansion)
Token IDs (int32)
0.40 MB
2.23 MB
0.18× (expansion)
GPT-2 hidden states
304.1 MB
2.23 MB
136.4× compression
v3.12 vs v3.7 Total Storage (extrapolated to 500 entries)
Version
Total Storage
vs UTF-8
vs Hidden States
v3.7
38.11 MB
0.32×
79.3×
v3.12
22.29 MB
0.54×
136.4×
Savings
15.82 MB (41.5%)
How v3.12 Achieves Better Compression Without Losing Accuracy
Function
v3.7 Implementation
v3.12 Implementation
Storage
Accuracy
Token matching
Store wte_centroid[768], cosine compare
forward_maxsim real-time compute
-3072 B/entry
Higher
Cross-domain filter
Cosine threshold
Expanded overlap gating (hard filter)
No extra storage
More precise
Weight adjustment
None
per_memory_forward_maxsim
No extra storage
New capability
Trade-off: v3.12 trades ~2ms extra computation per retrieval for 3072 bytes less storage per entry, while simultaneously improving retrieval accuracy by 54%.
4. Kakeya-like Compression (integrated)
The kakeya_codec.py module provides additional compression for the remaining semantic_emb[768] field (81% of v3.12 MemEntry).
Construction
Global PCA: R^768 → R^d_eff (retain 99% variance)
Temporal direction separation: coeff → (α scalar, perp vector)
Spherical K-means on perp directions → K segment centers (Kakeya skeleton)
Each memory encoded as (seg_id, α, t, sparse_residual)
Results (N=10, auto-threshold)
Metric
Value
Codec active
Yes
d_eff
7
K (segments)
8
Encode-decode cosine error
< 0.05 avg, < 0.1 max
Compression ratio
1.21×
All tests pass
48/48
Projected Compression at Scale
N
PCA only
Kakeya-like
Kakeya+int8
1K
72.1%
73.3%
75.7%
10K
73.1%
77.5%
81.2%
100K
65.9%
74.5%
81.1%
1M
56.9%
70.1%
80.2%
10M
56.7%
70.1%
80.2%
Kakeya-like compression becomes increasingly advantageous over PCA at N > 100K due to its local segment parameterization vs PCA's global basis.