The architecture must support:
- local-first operation
- privacy-sensitive screenshot processing
- fast search
- incremental indexing
- large screenshot collections
- replaceable OCR implementation
- future local AI inference
- minimal background resource usage
- safe recovery from interrupted indexing
React UI
|
v
Tauri IPC Commands / Events
|
v
Application Services
|
+-------------------+-------------------+
| | |
v v v
Indexing Search Settings
|
+---------+---------+---------+
| | | |
v v v v
Scanner Watcher OCR Thumbnail
|
v
Persistence
|
v
SQLite + FTS5
Phase 3.5E Hybrid Per-Line OCR Architecture (hybrid_v2):
Screenshot File
│
▼
OcrEngineRouter (Auto / Windows / Multilingual)
│
├── [Windows Mode] ──────────► WindowsMediaOcrEngine (WinRT native)
│
└── [Auto / Multilingual] ──► HybridOcrEngine
│
▼
DBNet Text Line Detector
│
▼
Text-line crops
│
▼
Windows OCR probe per crop
│
▼
Deterministic Line Classifier
(technical syntax first; structural Unicode
corruption signal for Windows Vietnamese)
│
┌───────────────┴───────────────┐
▼ ▼
Technical / Uncertain Natural Text
│ │
▼ ▼
KEEP Windows OCR VietOCR VGG-Transformer
│ │
└───────────────┬───────────────┘
▼
Reading-order merge
│
▼
Unicode NFC normalize
│
▼
OcrResult
│
┌─────────────────────────────┼─────────────────────────────┐
▼ ▼ ▼
SQLite `screenshots` SQLite FTS5 index Stale vector invalidate
(atomic metadata) (searchable text) (trigger semantic refresh)
Phase 3 Hybrid Search Architecture:
Query
|
+-----------------------------------+
| |
v v
SQLite FTS5 (Top 100) Local Embedding Model
(BM25 normalized) `multilingual-e5-small`
|
v
In-Process Cosine Scan
SQLite BLOB Vectors (Top 100)
|
+-----------------------------------+
|
v
Union Candidate Set
|
v
Hybrid Ranker
├── Exact Technical Token Guard (4.0x dominance)
├── Filename Matching (2.0x boost)
├── Normalized FTS Score (1.5x)
├── Normalized Semantic Score (1.2x)
└── Recency Tiebreak
|
v
Paginated Ranked Results (Top 50)
The frontend should be responsible for:
- rendering UI according to
docs/UI_DESIGN_SYSTEM.md - reusing shared shadcn-style primitives before creating custom controls
- accepting user input
- displaying indexing state
- search query state
- result rendering
- screenshot preview
- settings UI
- user-triggered actions
The frontend should NOT:
- invent feature-specific visual systems that conflict with
docs/UI_DESIGN_SYSTEM.md - duplicate existing shared UI primitives without justification
- recursively scan folders itself
- perform expensive OCR
- hash large files
- directly manipulate SQLite
- implement filesystem watchers
- run heavy AI inference on the UI thread
Rust is the trusted local core.
Suggested domains:
src-tauri/src/
├── commands/
├── db/
├── indexing/
├── ocr/
├── search/
├── thumbnails/
├── watcher/
├── settings/
├── filesystem/
└── errors/
Thin IPC boundary between frontend and native core.
Commands should:
- validate input
- invoke application services
- map errors to stable frontend-safe errors
Commands should not contain large amounts of business logic.
Responsibilities:
- connection initialization
- migrations
- repositories
- transactions
- FTS synchronization
- database health checks
Responsibilities:
- create indexing jobs
- coordinate scanner/OCR/thumbnail work
- enforce concurrency limits
- track progress
- retry recoverable failures
- skip unchanged files
Expose an OCR abstraction.
Conceptual interface:
OCRService
recognize(image_path) -> OCRResult
OCRResult should support at least:
- full normalized text
- raw text if needed internally
- optional text blocks
- optional bounding boxes
- detected language if available
- engine metadata/version
Indexing code must not depend directly on one OCR vendor.
Responsibilities:
- normalize query
- FTS query
- ranking
- filters
- result hydration
- future semantic search
- future hybrid ranking
Responsibilities:
- generate efficient preview images
- deterministic thumbnail paths
- invalidate thumbnails when source changes
- avoid loading original 4K images into search grid
Responsibilities:
- subscribe to selected folders
- debounce noisy file events
- detect create/update/delete
- schedule indexing
- avoid performing OCR directly inside watcher callback
User selects folder
|
v
Persist Folder
|
v
Recursive Scan
|
v
Supported File?
| |
no yes
| |
skip v
Read Metadata
|
v
File Identity
|
v
Existing unchanged?
| |
yes no
| |
skip v
Queue Job
|
v
Generate Thumb
|
v
OCR
|
v
Normalize Text
|
v
Commit DB
|
v
Update FTS
The exact ordering of thumbnail and OCR may be parallelized later.
Do not use filename alone.
A screenshot may be:
- renamed
- moved
- replaced
- edited
Possible identity inputs:
- canonical path
- size
- modification time
- hash
Recommended approach:
- use path for current location
- use content hash or a robust file fingerprint to determine whether content changed
- avoid hashing every large file repeatedly when metadata proves it is unchanged
The implementation may introduce a fast fingerprint and only calculate full hash when needed.
All heavy operations must be outside the UI thread.
Examples:
- recursive scanning
- hashing
- OCR
- thumbnail generation
- embedding generation
Use bounded concurrency.
Do not spawn unbounded work for thousands of images.
Desired behavior:
8,000 screenshots
|
v
bounded queue
|
+--> worker
+--> worker
+--> worker
|
v
steady progress
Concurrency should be configurable internally and later tunable based on CPU capabilities.
The backend should expose enough state for UI such as:
status
discovered_count
queued_count
processing_count
indexed_count
failed_count
skipped_count
total_count
Potential statuses:
IDLE
SCANNING
INDEXING
PAUSED
COMPLETED
CANCELLED
FAILED
Do not rely solely on frontend state for indexing truth.
Phase 1:
Query
|
v
Normalize
|
v
SQLite FTS5
|
v
BM25 ranking
|
v
Metadata filters
|
v
Hydrate results
|
v
UI
Future hybrid:
Query
|
+-------------------+
| |
v v
Keyword Embedding
| |
v v
FTS Score Semantic Score
| |
+---------+---------+
|
v
Hybrid Ranker
|
v
Results
Thumbnail cache should live in application-owned storage.
Example:
AppData/
└── ScreenshotSearch/
├── database.sqlite
├── thumbnails/
├── models/
└── logs/
Do not modify original screenshots.
Thumbnail filenames should derive from a stable identifier such as screenshot ID or content hash.
Failures should be categorized.
Examples:
FILE_NOT_FOUND
FILE_PERMISSION_DENIED
UNSUPPORTED_IMAGE
IMAGE_DECODE_FAILED
OCR_FAILED
DATABASE_FAILED
THUMBNAIL_FAILED
WATCHER_FAILED
MODEL_LOAD_FAILED
Do not collapse all failures into generic strings.
Recoverable failures should be retryable.
Permanent failures should not loop forever.
The app may close while indexing.
On restart:
- recover unfinished jobs
- verify source files still exist
- avoid duplicate OCR
- continue safely
- do not assume in-memory queue state survived
Initial MVP may use a simpler recoverable model, but architecture should allow persistent queue state later.
AI must remain behind dedicated services.
Conceptual interfaces:
TextEmbeddingService
VisualEmbeddingService
SemanticSearchService
Do not call AI models directly from React components.
Do not mix embedding model implementation with general database repositories.
When this document disagrees with actual code:
- inspect the implementation;
- determine whether the code or document is outdated;
- do not silently assume this document is correct;
- update the document after architectural changes are intentionally accepted.
Phase 4 adds an optional direct-pixel retrieval branch:
Screenshot pixels -> EXIF/RGBA decode -> letterbox + bounded tiles
-> CLIP image encoder -> SQLite visual vectors
Query -> multilingual CLIP-aligned text encoder -> visual query vector
FTS top K + text-semantic top K + visual top K -> union -> ranker
The pinned Qdrant CLIP ViT-B/32 vision encoder and sentence-transformers
multilingual CLIP text encoder produce normalized 512-dimensional vectors in the
same space. Normal screenshots use one global region; aspect ratios at least 1.8
add 2-6 tiles. Stored aggregation is 0.55 global + 0.45 tile mean; search uses
0.75 aggregate + 0.25 best region.
Visual work is independent of OCR and reuses the durable queue. A vector is valid only when content hash, model ID/version, and preprocessing version match. Exact/filename/FTS/text/visual rank weights are 10.0/2.0/1.5/1.2/1.3, with recency only as a tie-breaker. Separate image and text sessions allow queries to remain responsive during image backfill.