Summary
The chunks_fts virtual table uses the default FTS5 tokenizer, which splits text by Unicode word boundaries. For Chinese text, this means single-character tokenization — searching for "记忆系统" splits into four independent tokens "记", "忆", "系", "统", severely degrading BM25 precision.
Impact
Since the hybrid search weights BM25 at 30% of the final score, poor Chinese tokenization directly affects retrieval quality for any Chinese-language memory content. The vector search side (70%) partially compensates, but:
- BM25 is the fallback when embedding providers are unavailable (stub mode)
- BM25 excels at exact/precise matches that vector search sometimes misses
- Mixed Chinese/English content (common in this codebase) gets inconsistent treatment
Current Behavior
-- FTS5 default tokenizer splits Chinese by character:
-- "记忆系统设计" → ["记", "忆", "系", "统", "设", "计"]
-- Searching "记忆" matches any doc containing "记" OR "忆" individually
Suggested Approach
Options in order of complexity:
-
ICU tokenizer — FTS5 supports tokenize=icu which handles CJK segmentation via ICU library. Requires SQLite compiled with ICU support.
-
Custom tokenizer with jieba — Register a custom FTS5 tokenizer backed by jieba-rs for proper Chinese word segmentation. More accurate for Chinese but adds a dependency.
-
Bigram/trigram tokenizer — A simpler approach: generate overlapping n-grams for CJK characters. Less accurate than jieba but zero external dependencies.
References
crates/clawhive-memory/src/migrations.rs — migration 5: CREATE VIRTUAL TABLE IF NOT EXISTS chunks_fts USING fts5(text, ...)
crates/clawhive-memory/src/search_index.rs — BM25 search path
Summary
The
chunks_ftsvirtual table uses the default FTS5 tokenizer, which splits text by Unicode word boundaries. For Chinese text, this means single-character tokenization — searching for "记忆系统" splits into four independent tokens "记", "忆", "系", "统", severely degrading BM25 precision.Impact
Since the hybrid search weights BM25 at 30% of the final score, poor Chinese tokenization directly affects retrieval quality for any Chinese-language memory content. The vector search side (70%) partially compensates, but:
Current Behavior
Suggested Approach
Options in order of complexity:
ICU tokenizer — FTS5 supports
tokenize=icuwhich handles CJK segmentation via ICU library. Requires SQLite compiled with ICU support.Custom tokenizer with jieba — Register a custom FTS5 tokenizer backed by
jieba-rsfor proper Chinese word segmentation. More accurate for Chinese but adds a dependency.Bigram/trigram tokenizer — A simpler approach: generate overlapping n-grams for CJK characters. Less accurate than jieba but zero external dependencies.
References
crates/clawhive-memory/src/migrations.rs— migration 5:CREATE VIRTUAL TABLE IF NOT EXISTS chunks_fts USING fts5(text, ...)crates/clawhive-memory/src/search_index.rs— BM25 search path