Skip to content

FTS5 default tokenizer produces poor results for Chinese text #42

Description

@longzhi

Summary

The chunks_fts virtual table uses the default FTS5 tokenizer, which splits text by Unicode word boundaries. For Chinese text, this means single-character tokenization — searching for "记忆系统" splits into four independent tokens "记", "忆", "系", "统", severely degrading BM25 precision.

Impact

Since the hybrid search weights BM25 at 30% of the final score, poor Chinese tokenization directly affects retrieval quality for any Chinese-language memory content. The vector search side (70%) partially compensates, but:

  • BM25 is the fallback when embedding providers are unavailable (stub mode)
  • BM25 excels at exact/precise matches that vector search sometimes misses
  • Mixed Chinese/English content (common in this codebase) gets inconsistent treatment

Current Behavior

-- FTS5 default tokenizer splits Chinese by character:
-- "记忆系统设计" → ["记", "忆", "系", "统", "设", "计"]
-- Searching "记忆" matches any doc containing "记" OR "忆" individually

Suggested Approach

Options in order of complexity:

  1. ICU tokenizer — FTS5 supports tokenize=icu which handles CJK segmentation via ICU library. Requires SQLite compiled with ICU support.

  2. Custom tokenizer with jieba — Register a custom FTS5 tokenizer backed by jieba-rs for proper Chinese word segmentation. More accurate for Chinese but adds a dependency.

  3. Bigram/trigram tokenizer — A simpler approach: generate overlapping n-grams for CJK characters. Less accurate than jieba but zero external dependencies.

References

  • crates/clawhive-memory/src/migrations.rs — migration 5: CREATE VIRTUAL TABLE IF NOT EXISTS chunks_fts USING fts5(text, ...)
  • crates/clawhive-memory/src/search_index.rs — BM25 search path

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2: mediumImportant but not urgent, schedule itenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions