diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 690ef17..bfa6650 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -102,3 +102,16 @@ jobs: - name: Install cargo-audit run: cargo install cargo-audit --locked - run: cargo audit + + invariant: + name: Architecture Invariants + runs-on: windows-latest + steps: + - uses: actions/checkout@v4 + with: + fetch-depth: 0 + - name: Run invariant checks + shell: pwsh + run: | + git fetch origin main + tools/invariant-checks/run-checks.ps1 diff --git a/AGENTS.md b/AGENTS.md index ff93c06..9aa607f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -4,9 +4,9 @@ > 它将本地数字资产的原始数据(代码库、笔记、Skill、工作流)编译为 AI 可决策的结构化情境,不负责思考,不负责执行,只负责感知、编码、持久化、检索。 -- **当前阶段**:阶段六 — v0.14.3 / MCP Python SDK 兼容修复 + repo.rs trait 化收尾 -- **当前版本**:v0.14.3 (`main@2867811`) -- **已完成里程碑**:Registry God Object 完全拆解(10 子模块提取)+ 18 workspace crates 提取 + MCP Python SDK 1.16.0 兼容修复(NDJSON + null-id workaround)+ repo.rs crate:: 引用 13→9 + flaky 测试根治(RF-2.1/2.2/2.3)+ 许可证迁移(MIT → AGPL-3.0-or-later 双许可) +- **当前阶段**:阶段六 → v0.15.0 已发布(CHANGELOG 已记录);v0.16.0 进行中(Workspace 扩展 Phase 2) +- **当前版本**:v0.15.0(CHANGELOG 已记录,`fix/project-health-cleanup@4e8d882`,**待合入 `main` 并打 tag**) +- **已完成里程碑**:Registry God Object 完全拆解(10 子模块提取)+ 18 workspace crates 提取 + MCP Python SDK 1.16.0 兼容修复 + repo.rs trait 化 + flaky 测试根治(RF-2.1/2.2/2.3)+ 许可证迁移 + health 性能优化(-44%)+ index skip-embeddings + batch encoding 实验 + RF-6 清零 + 架构治理文档(ADR/不变量清单)+ Tantivy BM25 代码符号搜索(P1)+ AppContext 职责拆分 Phase 1/2(storage.rs 860→430 行)+ 架构不变量 CI(G5/T11/T12)+ Embedding 多后端(Candle/Ollama 配置切换, P3)+ EnvVersionCache 扩展(9 工具链检测, P4) - **核心方向**:让 Kimi CLI 在调用文件工具之前,先通过 devbase 获得"该读哪些文件、为什么读、它们之间的关系" - **本质分析**:见 `vault/99-Meta/devbase-essence-analysis-20260430.md` 与 `docs/architecture/redefinition.md` - **设计文档**: @@ -207,6 +207,24 @@ grep -rn "unwrap()\|expect()\|panic!(" src/ \ # 预期输出:空 ``` +### 架构治理框架(Architecture Governance) + +> 参考:外部架构治理方法论(Kimi 会话 `e9f2965f-b949-46a5-9d7c-afd6d4d9232c`) + +**已制度化实践**: + +| 实践 | devbase 落地形式 | 文档位置 | +|------|-----------------|---------| +| ADR(架构决策记录) | ADR-001(单 crate defer)、ADR-002(batch encoding 回滚) | [`docs/architecture/adr-template.md`](docs/architecture/adr-template.md) | +| 不变量清单(Invariants) | RF-1~RF-7 + 分层模块约束(T01–T12) | [`docs/architecture/invariants.md`](docs/architecture/invariants.md) | +| 模块提取演习 | RF-7 的 5 个 `crate::` 引用阈值 + 已提取 18 workspace crates | 本文件 §RF-7 | +| 三层摘要 | `crates/*/README.md` 要求:一句话 + 一页纸 + 深度链接 | 各 crate README | +| 定期架构回顾 | 每次 Wave 结束时的架构审计(见 `docs/_audit/`) | `docs/_audit/2026-04-26-*.md` | + +**待增强**: +- 三层摘要:部分已提取 crate 的 README 尚未达到"一页纸"标准 +- 定期架构回顾:当前按 Wave(功能迭代周期)触发,建议每 2–4 周增加一次纯架构 review(不看 feature 进度,只看不变量违反和隐式依赖) + --- ## 技术债登记簿(Technical Debt Ledger) @@ -216,13 +234,13 @@ grep -rn "unwrap()\|expect()\|panic!(" src/ \ | 债项 | 严重 | 当前值 | 目标阈值 | 清理路径 | 引入 Wave | |---|---|---|---|---|---| | `main.rs` 上帝文件 | 🟢 | 548 行 | ≤1000 行 | 拆分为 `commands/simple.rs` + `commands/skill.rs` + `commands/workflow.rs` + `commands/limit.rs`;全部 22 个命令/子命令树已迁移 | ≤15 | -| `init_db()` 全局路径 | 🟡 | `AppContext` 已集成到全部 commands/ 模块(22 个函数);`main()` 通过 `AppContext` 分发配置 | 0 新增 | `StorageBackend` trait + `AppContext` 已奠基;`db_path`/`workspace_dir`/`index_path`/`backup_dir` 已统一;`init_db()` 调用点 grandfathered 待迁移 | ≤15 | -| Tantivy+SQLite 双写一致性 | 🟡 | 无事务协调 | 补偿机制 | 设计 `sync_index_to_db()` 回滚或两阶段提交;或改为 SQLite FTS5 替代 Tantivy | 7 | +| `init_db()` 全局路径 | 🟢 | `AppContext` 已集成到全部 commands/ 模块;`main()` 通过 `AppContext` 分发配置;`init_db()` 无外部调用 | 0 | 已完成:`StorageBackend` trait + `AppContext` 全面替代;`db_path`/`workspace_dir`/`index_path`/`backup_dir` 已统一 | ≤15 | +| Tantivy+SQLite 双写一致性 | 🟡 | 无事务协调;**已添加反向检测**(`repair_tantivy_consistency` 现在检测 SQLite→Tantivy 缺失) | 补偿机制 | 长期:事务协调或 SQLite FTS5 替代;短期:反向检测 + 日志已落地(`fe14c81`) | 7 | | 主从表切换 | 🟢 | Phase 1 全部完成:`repos` 表已删除,entities 为唯一数据源 | `entities` 为第一公民 | Phase 2 类型系统开放(新增 entity_type 无需改表结构) | v0.12.0 | | vault/paper/workflow entities 缺口 | 🟢 | Stage C+D+E 全部完成:`vault_notes`/`papers`/`workflows` 表已删除,`skills` 保留(embedding BLOB) | 0 缺口 | — | v0.12.0 | | scan 路径排除 | 🟢 | `discover_repos` + `collect_tasks` 均支持 `scan.exclude_paths`;scan 和 sync 双阶段过滤 | 0 缺口 | 排除路径使用 `Path::starts_with` 组件级匹配,避免字符串前缀误杀;相对路径在 sync 场景(无 root)下被忽略 | v0.12.0 | -| tree-sitter 编译成本 | 🟡 | ~15-20s | 可控 | 评估 `ccache` 或 grammar 预编译;或按需 feature-gate | 8 | -| Feature flags 缺失 | 🟡 | 2 个可选 feature (tui, watch) | ≥3 (tui, watch, mcp) | `Cargo.toml` 已添加 `tui` + `watch` feature;ratatui/crossterm/notify 均为 optional;`--no-default-features` 编译通过 | ≤15 | +| tree-sitter 编译成本 | 🟢 | ~15-20s grammar C compilation | 可控 | 已完成 feature-gate:`lang-rust`/`lang-python`/`lang-js-ts`/`lang-go` 四个 feature,默认全启,可选关闭减少编译;`--no-default-features` 编译通过 | 8 | +| Feature flags 缺失 | 🟢 | 4 个可选 feature (tui, watch, mcp, embedding) | ≥3 | 已完成:`tui`/`watch`/`mcp`/`embedding` 均为 optional;`--no-default-features` 编译通过 | ≤15 | | `LOCALAPPDATA` 测试模式残留 | 🟢 | 0 处 | 0 | 全面废弃 `LOCALAPPDATA` 环境变量覆盖,统一为 `DEVBASE_DATA_DIR`;mcp/tests.rs 修复 cleanup 逻辑(remove_var 目标从 LOCALAPPDATA 修正为 DEVBASE_DATA_DIR) | 47 | **清偿原则**: @@ -329,7 +347,7 @@ devbase 承载外部资源调度的抽象接口: - **短期**:devbase MCP 接口可封装外部 TEE 服务(如 Azure Confidential Computing) - **长期**:如需自建,新建 `clarity-tee` 或 `devbase-secure` 子项目 -## 当前阶段待办(v0.12.0 发布冲刺) +## 当前阶段待办(v0.15.0 推进中) v0.11.3 已交付(tagged)。v0.12.0-alpha 全部功能已完成,进入发布治理阶段。 @@ -414,6 +432,15 @@ v0.11.3 已交付(tagged)。v0.12.0-alpha 全部功能已完成,进入发 - 生成工具:`tools/embedding-provider/skills.py`(sentence-transformers `all-MiniLM-L6-v2`) - 激活路径:启动 Ollama + `devbase index ` 生成 embedding,或配置远程 provider 于 `config.toml [embedding]` 段 +### 2026-05-04 索引性能实验记录 + +**发现**:Candle CPU BERT `batch_size=32` forward 比 `rayon` 并行单条慢 **5.2×**(88s vs 16s)。 +- 根因:Candle CPU matmul 对大 padded batch 不友好;batch 内序列长度差异导致大量无效 padding token 计算。 +- **决策**:`generate_and_save_embeddings` 回滚到 `rayon::par_iter()` 单条编码;保留 `EmbeddingProvider::encode_batch` trait 方法供未来 GPU/ONNX provider。 +- **新增**:`devbase index --skip-embeddings` 跳过 embedding 生成,纯符号/调用图索引从 ~16s 降至 ~250ms。 + +**外部参考**:知识蒸馏 Pipeline 设计规格(六阶段:噪声过滤→语义分割→主题聚类→层级展开→可信度标注→结构化输出),来源见 `docs/_audit/2026-04-26-embedding-research.md` §2026-05-04 补充。该规格提出通过 devbase MCP 暴露 `devkit_knowledge_distill` 工具,与 Vault 系统形成输入-处理-输出闭环。状态:设计规格级,待验证后评估集成优先级。 + ## 上下文安全机制(Context Safety Mechanism) > 长期架构原则:在多 Agent / 子代理协作场景下,保证工作区状态的一致性与可恢复性。 diff --git a/CHANGELOG.md b/CHANGELOG.md index eb4d772..3022bd9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,48 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [0.15.0] - 2026-05-04 + +### Added + +- **P1: Tantivy BM25 代码符号搜索** — `search/symbol_index.rs` + - 独立 Schema (`repo_id`, `name`, `signature`, `file_path`, `line_start`) + - `keyword_search_symbols` 主路径走 Tantivy BM25,SQLite LIKE 回退 + - 索引流程 `index.rs` 自动同步写入 symbol_index + - `StorageBackend` 扩展 `symbol_index_path()`(6 实现) +- **P3: Embedding 多后端** — Candle (默认) + Ollama (配置切换) + - 新增 `OllamaProvider` (`ureq` HTTP `/api/embed`) + - `create_provider(backend, model, base_url, timeout)` 配置化创建 + - `generate_query_embedding` 通过 `OnceLock` 懒加载配置化 provider + - 默认模型改为 `all-minilm` (384-dim,与 Candle 维度兼容) +- **P4: Health 环境检测扩展** — `EnvVersionCache` 从 5 工具 → 9 工具 + - 新增: `python`, `bun`, `zig`, `java` + - `get_tool_version` 支持 stderr fallback (Java 输出到 stderr) + - `fmt_version` 改进: Java 引号提取、Docker/Python 格式处理 +- **P5: 架构不变量自动化 CI** — `tools/invariant-checks/run-checks.ps1` + - G5: diff-only 检测新增生产代码 unwrap/expect/panic(排除 `#[cfg(test)]`) + - T11: 检测 `mcp/tools/*` 直接调用 `rusqlite::Connection` + - T12: 检测 `tui/render/*` 写入操作 + - CI job `invariant-check` 加入 `.github/workflows/ci.yml` +- **P2 Phase 1: AppContext 职责拆分** — 6 个 Client trait impl 迁出 `storage.rs` + - `scan.rs` / `health.rs` / `sync.rs` / `digest.rs` / `knowledge_engine/mod.rs` / `registry.rs` + - `storage.rs` 860 → 430 行 (-50%) + - 删除冗余 `conn_mut()` +- **P2 Phase 2: 内联 SQL 下沉** — 新增 `registry/code_symbols.rs` + `registry/dead_code.rs` + - `CodeSymbolRow` / `DeadCodeRow` + 纯函数查询 (12 个单元测试) + - `RegistryClient` 退化为纯代理层 + +### Changed + +- `EmbeddingConfig` 默认模型 `nomic-embed-text` → `all-minilm` (384-dim) +- AGENTS.md 阶段描述更新: v0.14.3 → v0.15.0 推进中 → v0.15.0 全部完成 + +### Fixed + +- **TTL 缓存负值 bug** (`97172ec`): `elapsed < ttl_seconds` → `elapsed >= 0 && elapsed < ttl_seconds` + - 防止系统时间回溯导致缓存永不过期 +- `crates/devbase-embedding/src/lib.rs` 遗留 unwrap 清零 (`encode_with_candle` → `ok_or`) + ## [0.14.3] - 2026-05-05 ### Added diff --git a/Cargo.lock b/Cargo.lock index 78e65fb..8f32452 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -1056,7 +1056,7 @@ checksum = "abd57806937c9cc163efc8ea3910e00a62e2aeb0b8119f1793a978088f8f6b04" [[package]] name = "devbase" -version = "0.14.3" +version = "0.15.0" dependencies = [ "anyhow", "assert_cmd", @@ -1134,6 +1134,7 @@ dependencies = [ "hf-hub", "serde_json", "tokenizers", + "ureq 2.12.1", ] [[package]] @@ -2091,7 +2092,7 @@ dependencies = [ "serde_json", "thiserror 2.0.18", "tokio", - "ureq", + "ureq 3.3.0", "windows-sys 0.61.2", ] @@ -5238,6 +5239,24 @@ version = "0.9.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1" +[[package]] +name = "ureq" +version = "2.12.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "02d1a66277ed75f640d608235660df48c8e3c19f3b4edb6a263315626cc3c01d" +dependencies = [ + "base64 0.22.1", + "flate2", + "log", + "once_cell", + "rustls", + "rustls-pki-types", + "serde", + "serde_json", + "url", + "webpki-roots 0.26.11", +] + [[package]] name = "ureq" version = "3.3.0" @@ -5259,7 +5278,7 @@ dependencies = [ "ureq-proto", "utf8-zero", "webpki-root-certs", - "webpki-roots", + "webpki-roots 1.0.7", ] [[package]] @@ -5534,6 +5553,15 @@ dependencies = [ "rustls-pki-types", ] +[[package]] +name = "webpki-roots" +version = "0.26.11" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "521bc38abb08001b01866da9f51eb7c5d647a19260e00054a8c7fd5f9e57f7a9" +dependencies = [ + "webpki-roots 1.0.7", +] + [[package]] name = "webpki-roots" version = "1.0.7" diff --git a/Cargo.toml b/Cargo.toml index e7da488..a568424 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "devbase" -version = "0.14.3" +version = "0.15.0" edition = "2024" description = "Developer workspace database and knowledge-base manager" authors = ["juice094 <160722440+juice094@users.noreply.github.com>"] @@ -14,11 +14,15 @@ name = "registry_bench" harness = false [features] -default = ["tui", "mcp", "embedding"] +default = ["tui", "mcp", "embedding", "lang-rust", "lang-python", "lang-js-ts", "lang-go"] tui = ["ratatui", "crossterm", "watch"] watch = ["notify"] mcp = [] embedding = ["dep:devbase-embedding"] +lang-rust = ["dep:tree-sitter-rust"] +lang-python = ["dep:tree-sitter-python"] +lang-js-ts = ["dep:tree-sitter-typescript"] +lang-go = ["dep:tree-sitter-go"] [dependencies] clap = { version = "4", features = ["derive"] } @@ -48,10 +52,10 @@ which = "7" dunce = "1" tokei = "14.0.0" tree-sitter = "0.26" -tree-sitter-rust = "0.24" -tree-sitter-python = "0.23" -tree-sitter-typescript = "0.23" -tree-sitter-go = "0.23" +tree-sitter-rust = { version = "0.24", optional = true } +tree-sitter-python = { version = "0.23", optional = true } +tree-sitter-typescript = { version = "0.23", optional = true } +tree-sitter-go = { version = "0.23", optional = true } serde_yaml = "0.9" regex = "1" # Extracted crates (zero internal coupling) @@ -89,7 +93,7 @@ members = ["crates/*"] resolver = "2" [workspace.package] -version = "0.14.3" +version = "0.15.0" authors = ["juice094 <160722440+juice094@users.noreply.github.com>"] edition = "2024" license = "AGPL-3.0-or-later" diff --git a/README.md b/README.md index ba856ad..e7c9379 100644 --- a/README.md +++ b/README.md @@ -1,7 +1,7 @@ # devbase -[![Version](https://img.shields.io/badge/version-v0.14.3-blue)](https://github.com/juice094/devbase/releases) -[![Tests](https://img.shields.io/badge/tests-390%20passed-brightgreen)](./AGENTS.md) +[![Version](https://img.shields.io/badge/version-v0.15.0-blue)](https://github.com/juice094/devbase/releases) +[![Tests](https://img.shields.io/badge/tests-490%2B%20passed-brightgreen)](./AGENTS.md) [![Clippy](https://img.shields.io/badge/clippy-0%20warnings-green)](./AGENTS.md) [![License](https://img.shields.io/badge/license-AGPL--3.0-orange)](./LICENSE) [![Rust](https://img.shields.io/badge/rust-1.94%2B-9cf)](https://www.rust-lang.org) @@ -283,19 +283,19 @@ TUI `[:]` 触发 embedding 语义搜索,失败自动降级为文本搜索。AI | v0.7.0 | ✅ 已发布 | NLQ 自然语言查询 + 智能同步建议 | | v0.8.0 | ✅ 已发布 | Workflow 子类型:Subworkflow / Parallel / Condition / Loop | | v0.9.0 | ✅ 已发布 | Loop Step 硬化 + 发布闭环 | -| **v0.10.0** | **✅ 已发布** | **L3-L4 知识模型 + 工程健康维护(main.rs 拆分 / StorageBackend / feature flags)** | -| **v0.11.0** | **✅ 已发布** | **AppContext Pool 化 + MCP 测试隔离 + CI 多线程** | -| **v0.11.1** | **✅ 已发布** | **Flat ID 命名空间 + entities-first 写入反转** | -| **v0.11.2** | **✅ 已发布** | **读路径全量迁移:所有 SELECT 切到 `entities`** | -| **v0.11.3** | **✅ 已发布** | **`repos` 表删除,`entities` 成为唯一数据源(Phase 1 完成)** | -| **v0.12.0-alpha** | **✅ 已发布** | **Phase 2 完成 (Stage A-E): entities 统一重构 + `.devbase-ignore` + managed-gate fail-safe 同步** | -| **v0.13.0** | **✅ 已发布** | **Registry God Object 拆解:10 子模块提取为 free function;WorkspaceRegistry 退化为纯 facade** | -| **v0.14.0** | **✅ 已发布** | **Workspace 拆分:6 个零耦合 crate 提取;MCP trait 化:`mcp/tools/repo.rs` `crate::` 引用 68→41** | -| **v0.15.0** | **✅ 已发布** | **Sprint A/B/C:三维 embedding + Saga 一致性 + Agent 状态接口 + MCP Streaming(45 tools)** | -| **v0.16.0** | **✅ 已发布** | **Workspace crate 第二批提取 + Windows debug 栈溢出修复** | -| **v0.14.1** | **✅ 已发布** | **CLI JSON 输出补全 (`--json`/`--recalc`) + relations MCP 工具加固 + License headers + Vault Daily/Graph** | -| **v0.14.2** | **✅ 已发布** | **health dirty 检测修复(排除 ignored 文件)+ scan 路径规范化 + syncthing-rust 识别修复 + experiment_log/CodeMetrics/ModuleGraph/CallGraph/DeadCode 提升为 Beta(48 tools: Stable 5 / Beta 40 / Experimental 3)** | -| **v0.14.3** | **✅ 当前** | **Schema v30 code symbol attributes + dead-code 过滤增强 + init_db() 注入式改造(RF-1)+ Tantivy/SQLite 补偿扫描 + Feature flags(mcp / embedding)+ sccache 构建加速文档** | +| v0.10.0 | ✅ 已发布 | L3-L4 知识模型 + 工程健康维护(main.rs 拆分 / StorageBackend / feature flags) | +| v0.11.0 | ✅ 已发布 | AppContext Pool 化 + MCP 测试隔离 + CI 多线程 | +| v0.11.1 | ✅ 已发布 | Flat ID 命名空间 + entities-first 写入反转 | +| v0.11.2 | ✅ 已发布 | 读路径全量迁移:所有 SELECT 切到 `entities` | +| v0.11.3 | ✅ 已发布 | `repos` 表删除,`entities` 成为唯一数据源(Phase 1 完成) | +| v0.12.0 | ✅ 已发布 | Phase 2 完成(Stage A-E):entities 统一重构 + `.devbase-ignore` + managed-gate fail-safe 同步 | +| v0.13.0 | ✅ 已发布 | Registry God Object 拆解:10 子模块提取为 free function;WorkspaceRegistry 退化为纯 facade | +| v0.14.0 | ✅ 已发布 | Workspace 拆分:6 个零耦合 crate 提取;MCP trait 化:`mcp/tools/repo.rs` `crate::` 引用 68→41 | +| v0.14.1 | ✅ 已发布 | CLI JSON 输出补全 (`--json`/`--recalc`) + relations MCP 工具加固 + License headers + Vault Daily/Graph | +| v0.14.2 | ✅ 已发布 | health dirty 检测修复(排除 ignored 文件)+ scan 路径规范化 + syncthing-rust 识别修复 + experiment_log/CodeMetrics/ModuleGraph/CallGraph/DeadCode 提升为 Beta(48 tools: Stable 5 / Beta 40 / Experimental 3) | +| v0.14.3 | ✅ 已发布 | Schema v30 code symbol attributes + dead-code 过滤增强 + init_db() 注入式改造(RF-1)+ Tantivy/SQLite 补偿扫描 + Feature flags(mcp / embedding)+ sccache 构建加速文档 | +| **v0.15.0** | **✅ 当前** | **P1 Tantivy BM25 代码符号搜索 + P2 AppContext 职责拆分(storage.rs 860→430 行)+ P3 Embedding 多后端(Candle + Ollama)+ P4 EnvVersionCache 扩展(9 工具链:含 python/bun/zig/java)+ P5 架构不变量自动化 CI(G5/T11/T12)** | +| v0.16.0 | 📋 进行中 | Workspace 扩展 Phase 2:成员目标 8-10 个;以 `docs/ROADMAP.md` 阶段六规划为准 | --- diff --git a/RELEASE_NOTES_v0.15.0.md b/RELEASE_NOTES_v0.15.0.md new file mode 100644 index 0000000..380b98b --- /dev/null +++ b/RELEASE_NOTES_v0.15.0.md @@ -0,0 +1,150 @@ +# devbase v0.15.0 Release Notes + +**Release Date**: 2026-05-04(CHANGELOG 收录日期;正式 git tag 在合入 main 后补打) +**Schema Version**: v30(无 schema 变更) +**Tests**: 490+ workspace passed / 0 failed / 4 ignored +**Branch**: `fix/project-health-cleanup` + +--- + +## Highlights + +本次发布以 **`plans/v0.15.0-directions.md` 调研结论**为路线图,按 P1~P5 优先级闭环交付五项工程能力,覆盖**搜索基础设施**、**架构纯度**、**离线适配**、**环境感知**与**CI 守卫**五个维度。 + +### P1 · Tantivy BM25 代码符号搜索 + +将 `hybrid.rs` 中的 SQLite `LIKE` 降级关键词路径替换为 **Tantivy BM25 索引**,实现真正意义上的代码符号级全文检索。 + +- 新增 `search/symbol_index.rs`:独立 Schema (`repo_id`, `name`, `signature`, `file_path`, `line_start`) +- `keyword_search_symbols` 主路径走 Tantivy BM25,SQLite LIKE 作为冷启动 fallback +- 索引流程 `index.rs` 在生成 code symbols 时**同步写入** symbol_index +- `StorageBackend` 扩展 `symbol_index_path()`(6 个 backend 实现全部覆盖) + +**收益**:大仓库符号查询从 `LIKE '%token%'` 全表扫描升级为倒排索引 + BM25 评分。 + +--- + +### P2 · AppContext 职责拆分 + +`storage.rs` 此前承载了 7 个 Client trait 的实现(~860 行),违反 SRP 且阻碍单元测试。 + +**Phase 1**:6 个 Client trait impl 迁出 `storage.rs` +- `scan.rs` / `health.rs` / `sync.rs` / `digest.rs` / `knowledge_engine/mod.rs` / `registry.rs` +- `storage.rs` 860 → 430 行(**-50%**) +- 删除冗余 `conn_mut()` + +**Phase 2**:内联 SQL 下沉 +- 新增 `registry/code_symbols.rs` + `registry/dead_code.rs` +- `CodeSymbolRow` / `DeadCodeRow` + 纯函数查询(12 个单元测试) +- `RegistryClient` 退化为纯代理层(pure facade) + +**收益**:消除了 devbase 最大的耦合黑洞之一;为后续 `devbase-mcp` 独立发布扫清障碍。 + +--- + +### P3 · Embedding 多后端(Candle + Ollama) + +打破 `CandleProvider` 单后端 + 首次运行强制联网下载模型的痛点。 + +- 新增 `OllamaProvider`(`ureq` HTTP `/api/embed`) +- `create_provider(backend, model, base_url, timeout)` 配置化创建 +- `generate_query_embedding` 通过 `OnceLock` 懒加载配置化 provider +- 默认模型改为 `all-minilm`(384-dim,与 Candle 维度兼容,方便后端切换) + +**配置示例**(`config.toml`): + +```toml +[embedding] +backend = "candle" # 或 "ollama" +model = "all-minilm" +base_url = "http://127.0.0.1:11434" # Ollama only +timeout_seconds = 30 +``` + +--- + +### P4 · Health 环境检测扩展 + +`EnvVersionCache` 从 **5 个工具链** 扩展到 **9 个**。 + +| 新增工具 | 检测命令 | 备注 | +|:---|:---|:---| +| `python` | `python --version` | 兼容 `python3` 回退 | +| `bun` | `bun --version` | | +| `zig` | `zig version` | | +| `java` | `java -version` | **stderr 输出**,新增 stderr fallback 解析 | + +- `get_tool_version` 支持 stderr fallback(Java 习惯打到 stderr) +- `fmt_version` 改进:Java 引号提取、Docker/Python 格式处理 + +**收益**:现代开发环境(多语言混栈、容器化、jvm 生态)首次在 `devbase health` 中获得完整版本快照。 + +--- + +### P5 · 架构不变量自动化 CI + +将 `docs/architecture/invariants.md` 中的人工审查规则转化为可在 PR 中自动 enforce 的脚本。 + +- 新增 `tools/invariant-checks/run-checks.ps1` + - **G5**:diff-only 检测新增生产代码 `unwrap`/`expect`/`panic!`(自动排除 `#[cfg(test)]` 块) + - **T11**:检测 `mcp/tools/*` 直接调用 `rusqlite::Connection`(违反 MCP trait 化方向) + - **T12**:检测 `tui/render/*` 写入操作(违反 TUI 纯消费者约束) +- CI job `invariant-check` 加入 `.github/workflows/ci.yml` + +**收益**:RF-6(生产 0 panic)和模块边界约束(T11/T12)从"代码审查依赖"升级为"CI 强制门控"。 + +--- + +## Changed + +- `EmbeddingConfig` 默认模型:`nomic-embed-text` → `all-minilm`(384-dim 统一,方便 Candle/Ollama 切换) +- `AGENTS.md` 阶段描述更新:v0.14.3 → v0.15.0 已发布 +- 主 crate `Cargo.toml` 与 `[workspace.package].version` bump 至 `0.15.0` + +--- + +## Fixed + +- **TTL 缓存负值 bug**(`97172ec`):`elapsed < ttl_seconds` → `elapsed >= 0 && elapsed < ttl_seconds` + - 防止系统时间回溯(NTP 调整、跨时区夏令时切换)导致缓存条目永不过期 +- `crates/devbase-embedding/src/lib.rs` 遗留 `unwrap` 清零(`encode_with_candle` → `ok_or` 显式错误传播) + +--- + +## Schema Changes + +**无 schema 变更**。Schema 保持 v30(v0.14.3 引入的 `code_symbols.attributes` 列)。 + +--- + +## Architecture Health + +| 维度 | v0.14.3 | v0.15.0 | +|:---|:---:|:---:| +| 生产 `unwrap()` 数 | 0 | 0 | +| `cargo clippy --all-targets -D warnings` | ✅ | ✅ | +| `cargo fmt --check` | ✅ | ✅ | +| Workspace crates | 18 | **19** | +| `storage.rs` 行数 | 860 | **430(-50%)** | +| `EnvVersionCache` 覆盖工具 | 5 | **9** | +| 自动化不变量检查 | 手工 | **CI 强制(G5/T11/T12)** | +| Embedding provider | Candle only | **Candle + Ollama 配置切换** | + +--- + +## Upgrade Notes + +- 无 breaking change;现有 SQLite registry / Tantivy 索引继续工作。 +- 首次运行 v0.15.0 会自动建立 `symbol_index` 子目录(位置由 `StorageBackend::symbol_index_path()` 决定)。 +- 如需切换到 Ollama embedding,参见上方 P3 章节的 `config.toml` 示例。 + +--- + +## Known Issues + +- `search::tests::test_search_repos` / `test_search_vault` 在多线程下偶发 flaky(单线程通过)。根因未定位,已记录在 `docs/_audit/project-status-snapshot-2026-04-29.md` §B2。 +- Tantivy + SQLite 双写仍无事务级一致性,依靠 v0.14.3 引入的反向补偿扫描兜底。 + +--- + +*本 Release Notes 由接管 Agent 于 2026-05-11 补写,弥补 v0.11.0~v0.14.x 期间 Release Notes 文件缺失的历史问题。后续版本应在发布同步生成。* diff --git a/benches/semantic_index.rs b/benches/semantic_index.rs index c14da9f..f6824ee 100644 --- a/benches/semantic_index.rs +++ b/benches/semantic_index.rs @@ -36,7 +36,7 @@ fn bench_index_repo_full(c: &mut Criterion) { group.bench_with_input(BenchmarkId::new("scale", label), &path, |b, p| { b.iter(|| { - let result = index_repo_full(p); + let result = index_repo_full(p).unwrap(); black_box(result); }); }); diff --git a/crates/devbase-embedding/Cargo.toml b/crates/devbase-embedding/Cargo.toml index d44d63b..4473d65 100644 --- a/crates/devbase-embedding/Cargo.toml +++ b/crates/devbase-embedding/Cargo.toml @@ -9,6 +9,7 @@ repository = "https://github.com/juice094/devbase" [dependencies] anyhow = "1" serde_json = "1" +ureq = { version = "2", features = ["json"] } candle-core = { version = "0.10" } candle-nn = { version = "0.10" } candle-transformers = { version = "0.10" } diff --git a/crates/devbase-embedding/README.md b/crates/devbase-embedding/README.md new file mode 100644 index 0000000..e37fd57 --- /dev/null +++ b/crates/devbase-embedding/README.md @@ -0,0 +1,12 @@ +# devbase-embedding + +Pure-local text embedding generation with pluggable backends: Candle (all-MiniLM-L6-v2) and Ollama. Zero external API dependencies. + +## Why use it + +For projects that need local embeddings without pulling in Python. Provides a unified `EmbeddingProvider` trait with configurable backend switching. + +## Alternatives + +- `text-embeddings-inference` (HuggingFace): more powerful but requires Docker/Python +- `fastembed-rs`: pure Rust but limited model selection diff --git a/crates/devbase-embedding/src/lib.rs b/crates/devbase-embedding/src/lib.rs index 28b1a05..70d8b11 100644 --- a/crates/devbase-embedding/src/lib.rs +++ b/crates/devbase-embedding/src/lib.rs @@ -15,6 +15,12 @@ pub trait EmbeddingProvider: Send + Sync { /// Generate an embedding for a single query string. fn encode(&self, text: &str) -> anyhow::Result>; + /// Generate embeddings for a batch of texts. + /// Default implementation falls back to sequential single encoding. + fn encode_batch(&self, texts: &[&str]) -> anyhow::Result>> { + texts.iter().map(|t| self.encode(t)).collect() + } + /// Provider name for diagnostics. fn name(&self) -> &'static str; } @@ -25,6 +31,100 @@ pub fn default_provider() -> Box { Box::new(CandleProvider) } +/// Create a provider from configuration parameters. +/// `backend`: "candle" | "ollama" +/// `model`: model name (for Ollama, e.g. "all-minilm") +/// `base_url`: Ollama base URL (e.g. "http://localhost:11434") +/// `timeout_seconds`: HTTP timeout for Ollama +pub fn create_provider( + backend: &str, + _model: &str, + base_url: &str, + timeout_seconds: u64, +) -> Box { + match backend { + "ollama" => Box::new(OllamaProvider::new(base_url, _model, timeout_seconds)), + _ => Box::new(CandleProvider), + } +} + +// --------------------------------------------------------------------------- +// OllamaProvider — local HTTP embedding via Ollama /api/embed +// --------------------------------------------------------------------------- + +pub struct OllamaProvider { + base_url: String, + model: String, + timeout_seconds: u64, +} + +impl OllamaProvider { + pub fn new(base_url: &str, model: &str, timeout_seconds: u64) -> Self { + Self { + base_url: base_url.trim_end_matches('/').to_string(), + model: model.to_string(), + timeout_seconds, + } + } + + fn embed_inner(&self, inputs: Vec<&str>) -> anyhow::Result>> { + let url = format!("{}/api/embed", self.base_url); + let body = if inputs.len() == 1 { + serde_json::json!({ + "model": self.model, + "input": inputs[0], + }) + } else { + serde_json::json!({ + "model": self.model, + "input": inputs, + }) + }; + + let resp: serde_json::Value = ureq::post(&url) + .set("Content-Type", "application/json") + .timeout(std::time::Duration::from_secs(self.timeout_seconds)) + .send_json(body) + .map_err(|e| anyhow::anyhow!("Ollama API request failed: {}", e))? + .into_json() + .map_err(|e| anyhow::anyhow!("Ollama API JSON parse error: {}", e))?; + + let embeddings = resp + .get("embeddings") + .and_then(|v| v.as_array()) + .ok_or_else(|| anyhow::anyhow!("Ollama response missing embeddings: {}", resp))?; + + let mut results = Vec::with_capacity(embeddings.len()); + for emb in embeddings { + let vec: Vec = emb + .as_array() + .ok_or_else(|| anyhow::anyhow!("invalid embedding array in Ollama response"))? + .iter() + .map(|v| v.as_f64().unwrap_or(0.0) as f32) + .collect(); + results.push(vec); + } + Ok(results) + } +} + +impl EmbeddingProvider for OllamaProvider { + fn encode(&self, text: &str) -> anyhow::Result> { + self.embed_inner(vec![text])? + .into_iter() + .next() + .ok_or_else(|| anyhow::anyhow!("empty embedding result from Ollama")) + } + + fn encode_batch(&self, texts: &[&str]) -> anyhow::Result>> { + self.embed_inner(texts.to_vec()) + } + + fn name(&self) -> &'static str { + "ollama" + } +} + // --------------------------------------------------------------------------- // CandleProvider — pure-Rust local embedding via all-MiniLM-L6-v2 // --------------------------------------------------------------------------- @@ -36,6 +136,10 @@ impl EmbeddingProvider for CandleProvider { let (model, tokenizer) = get_candle_resources()?; encode_with_candle(model, tokenizer, text) } + fn encode_batch(&self, texts: &[&str]) -> anyhow::Result>> { + let (model, tokenizer) = get_candle_resources()?; + encode_batch_with_candle(model, tokenizer, texts) + } fn name(&self) -> &'static str { "candle-all-MiniLM-L6-v2" } @@ -87,29 +191,65 @@ fn encode_with_candle( tokenizer: &tokenizers::Tokenizer, text: &str, ) -> anyhow::Result> { + encode_batch_with_candle(model, tokenizer, &[text]) + .and_then(|mut v| v.pop().ok_or_else(|| anyhow::anyhow!("empty embedding batch"))) +} + +fn encode_batch_with_candle( + model: &candle_transformers::models::bert::BertModel, + tokenizer: &tokenizers::Tokenizer, + texts: &[&str], +) -> anyhow::Result>> { use candle_core::Tensor; + if texts.is_empty() { + return Ok(Vec::new()); + } - let encoding = tokenizer.encode(text, true).map_err(|e| anyhow::anyhow!(e))?; - let input_ids = encoding.get_ids(); - let attention_mask = encoding.get_attention_mask(); + // Batch tokenize + let encodings = tokenizer.encode_batch(texts.to_vec(), true).map_err(|e| anyhow::anyhow!(e))?; + + // Find max length for padding + let max_len = encodings.iter().map(|e| e.get_ids().len()).max().unwrap_or(0); + + // Build padded batch tensors + let mut input_ids_vec = Vec::new(); + let mut attention_mask_vec = Vec::new(); + for encoding in &encodings { + let ids = encoding.get_ids(); + let mask = encoding.get_attention_mask(); + let mut padded_ids = ids.to_vec(); + let mut padded_mask = mask.to_vec(); + padded_ids.resize(max_len, 0); + padded_mask.resize(max_len, 0); + input_ids_vec.extend(padded_ids); + attention_mask_vec.extend(padded_mask); + } - let input_ids = Tensor::new(input_ids, &model.device)?.unsqueeze(0)?; + let batch_size = texts.len(); + let input_ids = Tensor::new(input_ids_vec, &model.device)?.reshape((batch_size, max_len))?; let token_type_ids = input_ids.zeros_like()?; - let attention_mask_t = Tensor::new(attention_mask, &model.device)?.unsqueeze(0)?; + let attention_mask_t = + Tensor::new(attention_mask_vec, &model.device)?.reshape((batch_size, max_len))?; + // Single forward pass for the whole batch let output = model.forward(&input_ids, &token_type_ids, Some(&attention_mask_t))?; - // Mean pooling: average over non-padding tokens + // Mean pooling + L2 normalize per sample let mask = attention_mask_t.to_dtype(candle_core::DType::F32)?.unsqueeze(2)?; let sum = output.broadcast_mul(&mask)?.sum(1)?; let count = mask.sum(1)?; let mean_pooled = sum.broadcast_div(&count)?; - // L2 normalize (sentence-transformers default) let norm = mean_pooled.sqr()?.sum_keepdim(1)?.sqrt()?; let normalized = mean_pooled.broadcast_div(&norm)?; - Ok(normalized.squeeze(0)?.to_vec1()?) + // Extract per-sample embeddings + let mut results = Vec::with_capacity(batch_size); + for i in 0..batch_size { + let emb = normalized.get(i)?.squeeze(0)?.to_vec1()?; + results.push(emb); + } + Ok(results) } /// Cosine similarity between two f32 vectors. @@ -143,12 +283,6 @@ pub fn bytes_to_embedding(bytes: &[u8]) -> Vec { .collect() } -/// Convenience wrapper that uses the default provider. -/// Kept for backward compatibility with existing call sites. -pub fn generate_query_embedding(query: &str) -> anyhow::Result> { - default_provider().encode(query) -} - #[cfg(test)] mod tests { use super::*; @@ -273,4 +407,22 @@ mod tests { last_err )) } + + #[test] + fn test_create_provider_candle() { + let provider = create_provider("candle", "", "", 0); + assert_eq!(provider.name(), "candle-all-MiniLM-L6-v2"); + } + + #[test] + fn test_create_provider_ollama() { + let provider = create_provider("ollama", "all-minilm", "http://localhost:11434", 30); + assert_eq!(provider.name(), "ollama"); + } + + #[test] + fn test_create_provider_unknown_defaults_to_candle() { + let provider = create_provider("unknown", "", "", 0); + assert_eq!(provider.name(), "candle-all-MiniLM-L6-v2"); + } } diff --git a/crates/devbase-registry-health/src/lib.rs b/crates/devbase-registry-health/src/lib.rs index 0dc825c..62d6fe4 100644 --- a/crates/devbase-registry-health/src/lib.rs +++ b/crates/devbase-registry-health/src/lib.rs @@ -56,6 +56,50 @@ pub fn get_health( } } +/// Batch load health entries for multiple repos in a single query. +pub fn get_health_batch( + conn: &rusqlite::Connection, + repo_ids: &[&str], +) -> anyhow::Result> { + if repo_ids.is_empty() { + return Ok(std::collections::HashMap::new()); + } + let placeholders: Vec = repo_ids.iter().map(|_| "?".to_string()).collect(); + let sql = format!( + "SELECT repo_id, status, ahead, behind, checked_at FROM repo_health WHERE repo_id IN ({})", + placeholders.join(",") + ); + let mut stmt = conn.prepare(&sql)?; + let params: Vec<&dyn rusqlite::ToSql> = + repo_ids.iter().map(|id| id as &dyn rusqlite::ToSql).collect(); + let rows = stmt.query_map(rusqlite::params_from_iter(params.iter()), |row| { + let repo_id: String = row.get(0)?; + let status: String = row.get(1)?; + let ahead: i64 = row.get(2)?; + let behind: i64 = row.get(3)?; + let checked_at: String = row.get(4)?; + let checked_at = match DateTime::parse_from_rfc3339(&checked_at) { + Ok(dt) => dt.with_timezone(&Utc), + Err(_) => Utc::now(), + }; + Ok(( + repo_id, + HealthEntry { + status, + ahead: ahead as usize, + behind: behind as usize, + checked_at, + }, + )) + })?; + let mut result = std::collections::HashMap::new(); + for row in rows { + let (id, entry) = row?; + result.insert(id, entry); + } + Ok(result) +} + pub fn save_stars_cache( conn: &rusqlite::Connection, repo_id: &str, diff --git a/docs/README.md b/docs/README.md index 0f11f08..a378b0a 100644 --- a/docs/README.md +++ b/docs/README.md @@ -1,8 +1,8 @@ # devbase 文档导航 -> **项目状态**:v0.13.0 — 情境编译器闭环构建中 +> **项目状态**:v0.14.3 — 性能优化 + 架构治理 + RF-6 清零 > **主入口**:[`AGENTS.md`](../AGENTS.md)(Agent 环境指引)· [`ROADMAP.md`](ROADMAP.md)(功能路线图) -> **最后整理**:2026-04-30 +> **最后整理**:2026-05-08 --- @@ -10,12 +10,14 @@ | 指标 | 数值 | |------|------| -| 版本 | v0.13.0 | -| 测试 | 389 passed / 0 failed / 4 ignored | +| 版本 | v0.14.3 | +| 测试 | 408 passed / 0 failed / 3 ignored | | Clippy | 0 warnings | | Schema | v23 | | MCP Tools | 38 个(Stable 5 / Beta 29 / Experimental 4) | -| 代码行数 | ~30 KLOC | +| 代码行数 | ~32 KLOC | +| RF-6 | ✅ 生产代码 unwrap/expect 清零 | +| 已提取 Crates | 18 个 workspace crate | --- @@ -47,6 +49,8 @@ | [`architecture/workflow-dsl.md`](architecture/workflow-dsl.md) | Workflow DSL v0.4.0 规范(YAML 多步骤编排) | | [`architecture/dependency-topology.md`](architecture/dependency-topology.md) | 模块依赖拓扑(Tier 1–11 自底向上进化顺序) | | [`architecture/pre-split-evaluation.md`](architecture/pre-split-evaluation.md) | 单 crate vs 多 crate 评估结论 | +| [`architecture/adr-template.md`](architecture/adr-template.md) | ADR 模板与已完成决策索引(batch encoding、split defer) | +| [`architecture/invariants.md`](architecture/invariants.md) | 架构不变量清单:全局 + 分层约束 + 提取演习检查表 | ### 📖 使用指南(Guides) diff --git a/docs/_audit/2026-04-26-embedding-research.md b/docs/_audit/2026-04-26-embedding-research.md index d1ef9a7..9c91253 100644 --- a/docs/_audit/2026-04-26-embedding-research.md +++ b/docs/_audit/2026-04-26-embedding-research.md @@ -144,3 +144,48 @@ fn encode(model: &BertModel, tokenizer: &Tokenizer, text: &str) -> anyhow::Resul 3. 修改 `src/embedding.rs` — feature-gated 路由 (`local-embedding` 优先,否则 Python 回退) 4. 集成测试: 验证 embedding 维度 = 384,与 Python 输出余弦相似度 > 0.999 5. 更新 `project_context` — `goal` 参数启用时自动生成 query embedding + +--- + +## 2026-05-04 补充:batch encoding 实验与知识蒸馏参考 + +### batch encoding 实验(Candle CPU BERT) + +**假设**:batch_size=32 可提升 BERT 吞吐量,减少 1572 次单条 forward 的 overhead。 +**结果**:batch encoding **比单条慢 5.2 倍**(88s vs 16s)。 + +| 方案 | embed 时间 | total 时间 | symbols | +|------|-----------|-----------|---------| +| 单条并行 (`rayon::par_iter()`) | 15,941ms | 16,322ms | 1573 | +| batch=32 (`encode_batch_with_candle`) | 88,259ms | 88,987ms | 1573 | +| `--skip-embeddings` | 0ms | 245ms | 1573 | + +**根因分析**: +- Candle CPU 矩阵乘法对大 padded batch 不友好 +- batch 内序列长度差异导致 pad 到 max_len,处理大量无效 token +- 每 batch forward ~1.7s,而单条仅 ~10ms;49 batches × 1.7s >> 1572 × 10ms + +**决策**: +- `generate_and_save_embeddings` 回滚到 `rayon::par_iter()` 单条编码 +- 保留 `EmbeddingProvider::encode_batch` trait 方法供未来 GPU/ONNX provider +- 新增 `--skip-embeddings` CLI 标志:纯符号/调用图索引跳过 embedding(16s → 0.25s) + +### 外部文献参考:知识蒸馏 Pipeline + +> 来源:Kimi 会话研究(`kimi.com/share/19e06760-97a2-84ea-8000-00006aaf036d`) + +**`distill-knowledge` Skill 设计规格**(六阶段 Pipeline): +1. **噪声过滤**:去除 UI 元素、元数据噪音、重复路径、广告语句、时间戳 +2. **语义块分割**:切分为最小知识单元(项目/概念/论证) +3. **主题聚类**:一级主题 ≤6,二级分类 ≤4(奥卡姆剃刀) +4. **层级展开**:三级知识点表格化(项目|核心特征|可信度) +5. **可信度标注**:🟡 文档验证 / ⚠️ 社区传播 / ❓ 设想级 +6. **结构化输出**:Markdown 知识库(目录锚点/统计区块/跨域规律附录) + +**与 devbase 的对接点**: +- Vault 笔记系统可作为知识蒸馏的**输出存储**(`devkit_vault_write`) +- `devkit_paper_index`(PDF)可扩展为通用文本知识蒸馏入口 +- MCP Server 可暴露 `devkit_knowledge_distill` 工具,供 Clarity 等 Agent 调用做预处理 +- 噪声过滤规则库可直接嵌入 vault 索引流程,提升笔记质量 + +**状态**:设计规格级,待 Clarity 侧落地验证后评估 devbase 集成优先级。 diff --git a/PROJECT_STATUS_2026-04-29.md b/docs/_audit/project-status-snapshot-2026-04-29.md similarity index 100% rename from PROJECT_STATUS_2026-04-29.md rename to docs/_audit/project-status-snapshot-2026-04-29.md diff --git a/STAGE_REPORT_2026-04-09.md b/docs/_audit/stage-report-2026-04-09.md similarity index 100% rename from STAGE_REPORT_2026-04-09.md rename to docs/_audit/stage-report-2026-04-09.md diff --git a/docs/architecture/adr-template.md b/docs/architecture/adr-template.md new file mode 100644 index 0000000..a10634e --- /dev/null +++ b/docs/architecture/adr-template.md @@ -0,0 +1,82 @@ +# Architecture Decision Record (ADR) 模板 + +> 来源:架构治理方法论参考(Kimi 会话 `e9f2965f-b949-46a5-9d7c-afd6d4d9232c`) +> 用途:记录任何"为什么不那样做"的决策,一句话 + 一句话后果。 + +--- + +## ADR-XXX: [标题] + +- **状态**: proposed / accepted / deprecated / superseded by ADR-YYY +- **日期**: YYYY-MM-DD +- **作者**: [名字/会话ID] + +### 上下文 + +[背景:当时面临什么问题,有哪些可选方案] + +### 决策 + +[明确陈述:选择了什么] + +### 后果 + +- **正面**:[带来了什么好处] +- **负面**:[付出了什么代价] +- **风险**:[未来可能产生的问题] + +### 备选方案 + +| 方案 | 不选原因 | +|------|---------| +| [方案A] | [一句话理由] | +| [方案B] | [一句话理由] | + +### 相关决策 + +- 依赖:ADR-ZZZ +- 被依赖:ADR-WWW +- 替代:ADR-YYY + +--- + +## 已完成的 ADR 索引 + +| 编号 | 标题 | 状态 | 日期 | +|------|------|------|------| +| ADR-001 | 单 crate 模型(defer split)| accepted | 2026-04-26 | +| ADR-002 | Candle CPU BERT 单条编码(batch 回滚)| accepted | 2026-05-04 | + +### ADR-001: 单 crate 模型(defer split) + +- **状态**: accepted +- **日期**: 2026-04-26 +- **作者**: devbase 架构审计 + +**上下文**: v0.2.4 时 22.7 KLOC,评估是否拆分为 workspace crates。 + +**决策**: Defer split。保持单 crate,触发条件:50+ MCP tools / clean build > 60s / binary > 20 MB。 + +**后果**: +- 正面:迭代更快,无跨 crate 版本协调 +- 负面:编译缓存粒度粗,模块化约束靠自觉 +- 风险:长期可能积累隐式耦合,需定期提取演习验证 + +### ADR-002: Candle CPU BERT 单条编码(batch 回滚) + +- **状态**: accepted +- **日期**: 2026-05-04 +- **作者**: devbase 性能优化会话 + +**上下文**: Index 冷启动 17s,embedding 占 98.5%。尝试 batch_size=32 降低 forward 次数。 + +**决策**: 回滚到 `rayon::par_iter()` 单条编码;保留 `encode_batch` trait 供未来 GPU/ONNX provider。 + +**后果**: +- 正面:恢复 16s 基线;新增 `--skip-embeddings` 路径(0.25s) +- 负面:Candle CPU 上无法利用 batch 加速 +- 风险:未来切换 GPU provider 时需重新验证 batch 策略 + +--- + +*本文档遵循"一句话 + 一句话后果"原则,禁止长篇实现细节。* diff --git a/docs/architecture/invariants.md b/docs/architecture/invariants.md new file mode 100644 index 0000000..6803f9c --- /dev/null +++ b/docs/architecture/invariants.md @@ -0,0 +1,75 @@ +# 架构不变量清单(Invariants) + +> 来源:架构治理方法论参考(Kimi 会话 `e9f2965f-b949-46a5-9d7c-afd6d4d9232c`) +> 原则:不可打破的规则列表,每次代码审查对照检查。 + +--- + +## 全局不变量 + +| # | 规则 | 违反后果 | 检测方式 | +|---|------|---------|---------| +| G1 | `registry::WorkspaceRegistry` 不得依赖任何 Tier 4+ 模块 | 数据层被查询层污染,Schema 变更引发级联修改 | `cargo check` + 依赖拓扑审查 | +| G2 | `i18n` / `config` 不得包含业务逻辑 | 基础配置层膨胀,语言文件与业务耦合 | 代码审查:只含静态字符串和结构体定义 | +| G3 | 所有状态变更 MCP tool 必须幂等 | 重复调用导致数据损坏 | 单元测试:同一参数调用两次结果一致 | +| G4 | Breaking change 只能通过新增 tool 实现,不修改现有 schema | 下游 Agent 契约破裂 | Schema 版本对比 + MCP tool 清单审计 | +| G5 | 生产代码不得新增 `unwrap()` / `expect()`(RF-6) | 运行时 panic | `cargo clippy` + 人工审查 | + +## 分层不变量 + +### Tier 0–1(原子基础层) + +| # | 规则 | 说明 | +|---|------|------| +| T01 | `core` 只定义无业务语义的枚举和结构体 | NodeType / Node / Edge 不得出现 devbase 专属逻辑 | +| T02 | `registry` Schema 变更必须经过三步:migration → 备份 → oplog_analytics 兼容性检查 | 见 `dependency-topology.md` §二、Tier 1 | +| T03 | `embedding` 必须是纯函数工具包,无副作用 | 禁止在 encode 中写文件、改全局状态 | + +### Tier 2–3(扫描与分析层) + +| # | 规则 | 说明 | +|---|------|------| +| T04 | `scan` 新增语言支持不得改动 `semantic_index` 公共 API | 语言检测规则可独立实验 | +| T05 | `symbol_links` 的阈值和算法可独立调优,不破坏下游 | Jaccard 阈值默认 0.3,可调 | + +### Tier 4(查询层) + +| # | 规则 | 说明 | +|---|------|------| +| T06 | `query` 表达式解析必须向后兼容 | `lang:rust` 语法不得删除,只能扩展 | +| T07 | `search/hybrid` RRF 权重可独立调优,不影响工具 schema | 向量/BM25 融合策略是内部实现细节 | + +### Tier 5(同步层) + +| # | 规则 | 说明 | +|---|------|------| +| T08 | 新增 sync 策略必须先定义于 `sync/policy`,再实现于 `sync/tasks` | 禁止直接在 orchestrator 中硬编码策略逻辑 | + +### Tier 6–7(Skill / Workflow 层) + +| # | 规则 | 说明 | +|---|------|------| +| T09 | `skill_runtime::executor` 必须自包含副作用描述 | 每个 Skill 的 entry_script 必须声明读写范围 | +| T10 | Workflow 新增 `StepType` 只需改动 `workflow/model` → `parser` → `executor`,不影响 Skill Runtime | 见 `dependency-topology.md` §二、Tier 7 | + +### Tier 9–10(MCP / TUI 层) + +| # | 规则 | 说明 | +|---|------|------| +| T11 | `mcp/tools/*` 不得直接调用 `rusqlite::Connection`,必须通过 `registry` 封装 | 防止 SQL 注入和 Schema 漂移 | +| T12 | `tui/render/*` 是纯消费者层,新增面板不改动任何下层逻辑 | TUI 状态机只读取,不写入 registry | + +## 模块提取演习检查表 + +> 每季度执行一次。任何模块若无法在半天内提取为独立包并写出 50 字 README,说明耦合过重。 + +| 模块 | 上次检查 | 能否提取 | README 50 字验证 | +|------|---------|---------|-----------------| +| `devbase-embedding` | 2026-04 | ✅ 已提取 | 本地 BERT embedding 生成,零外部依赖 | +| `devbase-registry-health` | 2026-04 | ✅ 已提取 | 批量健康查询,消除 N+1 | +| `skill_runtime` | 2026-04 | ⚠️ 待验证 | parser/registry/executor 边界清晰,但 clarity_sync 耦合待审 | +| `workflow` | 2026-04 | ⚠️ 待验证 | model/parser/scheduler 可拆,executor 依赖 skill_runtime::executor | + +--- + +*违反任何不变量 → 必须在 ADR 中记录例外理由,否则回滚。* diff --git a/plans/appcontext-refactor-design.md b/plans/appcontext-refactor-design.md new file mode 100644 index 0000000..de937de --- /dev/null +++ b/plans/appcontext-refactor-design.md @@ -0,0 +1,212 @@ +# devbase AppContext 职责拆分设计文档 + +> **版本**: v0.15.0-P2 +> **日期**: 2026-05-08 +> **范围**: `src/storage.rs`、`src/clients.rs` 及 6 个 Client trait 实现 +> **目标**: 消除 `AppContext` 违反 SRP 的集中式 trait 实现,降低 `storage.rs` 耦合度,提升单元测试可行性。 + +--- + +## 一、现状诊断 + +### 1.1 规模与耦合 + +`src/storage.rs` 当前 **860 行**,其中约 **432 行**(50%)为 6 个 Client trait 的 `impl` 块: + +| Trait | 方法数 | 所在行号 | 依赖的核心模块 | +|-------|--------|----------|----------------| +| `ScanClient` | 1 | 309–317 | `scan::run_json` | +| `HealthClient` | 1 | 319–341 | `health::run_json`、`health::refresh_env_cache` | +| `SyncClient` | 1 | 343–353 | `sync::run_json` | +| `DigestClient` | 1 | 355–361 | `digest::generate_daily_digest` | +| `KnowledgeClient` | 4 | 363–408 | `knowledge_engine::run_index`、`registry::knowledge::*` | +| `RegistryClient` | 11 | 410–741 | `registry::repo`、`registry::knowledge`、`registry::metrics`、`registry::health`、`dependency_graph`、`registry::call_graph` | + +这些 `impl` 块与 `AppContext` 本体(字段定义、构造函数、连接池管理、`EnvVersionCache` 管理)共存于同一文件,导致: + +1. **单一文件职责过载**:`storage.rs` 同时承担"存储后端抽象"、"应用上下文生命周期管理"、"6 个业务领域 Client 实现"。 +2. **变更放大效应**:修改 `RegistryClient` 的查询逻辑(如 `query_code_symbols`)会触发 `storage.rs` 的重新编译,间接影响所有依赖 `StorageBackend` 的模块。 +3. **单元测试阻碍**:`AppContext` 的 `pool` 和 `env_cache` 为私有字段,测试只能通过 `with_storage()` 构造完整上下文;而业务逻辑与上下文构造代码混在一起,导致 mock 成本过高。 +4. **内联 SQL 泄漏**:`RegistryClient` 的 `query_code_symbols` 和 `query_dead_code` 直接在 `storage.rs` 中拼接 SQL,绕过了 `registry` 子模块的封装边界。 + +### 1.2 当前 AppContext 方法清单(按职责分组) + +#### A. 基础设施生命周期(应保留在 `storage.rs`) +- `with_defaults() -> Result` +- `with_storage(storage) -> Result` +- `build_pool(path) -> Result>` (private) +- `conn() -> Result>` +- `conn_mut() -> Result>` (与 `conn()` 完全等价,API 冗余) +- `pool() -> Pool<...>` +- `env_cache() -> Result` +- `set_env_cache(cache) -> Result<()>` + +#### B. 扫描职责 (`ScanClient`) +- `scan_directory(path, register) -> Future>` + +#### C. 健康职责 (`HealthClient`) +- `check_health(detail) -> Future>`(含 `env_cache` 刷新逻辑) + +#### D. 同步职责 (`SyncClient`) +- `sync_repos(dry_run, filter_tags) -> Future>` + +#### E. 摘要职责 (`DigestClient`) +- `generate_daily_digest() -> Result` + +#### F. 知识引擎职责 (`KnowledgeClient`) +- `run_index(path) -> Result` +- `save_note(repo_id, text, author) -> Result` +- `save_summary(repo_id, desc, author) -> Result` +- `get_paper(arxiv_id) -> Result` + +#### G. 注册表职责 (`RegistryClient`) +- `list_repos(filter) -> Result` +- `get_repo(repo_id) -> Result` +- `list_modules(repo_id) -> Result` +- `save_paper(paper) -> Result` +- `save_experiment(exp) -> Result` +- `list_code_metrics() -> Result` +- `get_code_metrics(repo_id) -> Result` +- `get_health(repo_id) -> Result` +- `query_call_graph(repo_id, callee, caller, file, limit) -> Result` +- `query_dependencies(repo_id, direction, relation_type) -> Result` +- `query_code_symbols(repo_id, name, symbol_type, file, limit) -> Result` +- `query_dead_code(repo_id, include_pub, limit) -> Result` + +--- + +## 二、拆分方案 + +### 2.1 核心原则 + +1. **AppContext 保持为依赖容器**:不拆分字段,不引入 6 个新的 Service 结构体以避免调用方大面积重构。 +2. **`impl` 块按职责归属迁移**:将 `impl XxxClient for AppContext` 从 `storage.rs` 剪切到对应的业务模块文件中。Rust 允许在任意模块中为外部类型实现外部 trait(orphan rules 在此不适用,因为 `AppContext` 和 trait 均定义在当前 crate)。 +3. **trait 定义不动**:`src/clients.rs` 保持为 MCP tool 与业务模块之间的稳定契约。 +4. **分阶段实施**:Phase 1 为纯物理迁移(零行为变更),Phase 2 为逻辑下沉(提取内联 SQL 到 `registry` 子模块)。 + +### 2.2 Phase 1:文件级 `impl` 块迁移(零行为变更) + +| Trait | 原位置 | 目标位置 | 理由 | +|-------|--------|----------|------| +| `ScanClient` | `storage.rs:309` | `scan.rs` | 直接代理 `scan::run_json`,归属扫描模块。 | +| `HealthClient` | `storage.rs:319` | `health.rs` | 直接代理 `health::run_json` 并管理 `env_cache`,归属健康模块。 | +| `SyncClient` | `storage.rs:343` | `sync.rs` | 直接代理 `sync::run_json`,归属同步模块。 | +| `DigestClient` | `storage.rs:355` | `digest.rs` | 直接代理 `digest::generate_daily_digest`,归属摘要模块。 | +| `KnowledgeClient` | `storage.rs:363` | `knowledge_engine/mod.rs` | 混合调用 `knowledge_engine::run_index` 与 `registry::knowledge`,归属知识引擎模块。 | +| `RegistryClient` | `storage.rs:410` | `registry.rs` | 大量调用 `registry::*` 子模块,归属注册表模块。 | + +**迁移后 `storage.rs` 的预期规模**: + +- `EnvVersionCache` + `StorageBackend` / `DefaultStorageBackend` / `TempStorageBackend`:~105 行 +- `AppContext` 本体(字段 + 构造函数 + `conn/pool/env_cache` 方法):~105 行 +- `repair_tantivy_consistency` / `repair_tantivy_consistency_at`:~90 行 +- `#[cfg(test)]` 测试:~75 行 +- **总计约 375 行**,较当前 **860 行** 减少 **56%**。 + +### 2.3 Phase 2:逻辑下沉与封装修复 + +Phase 1 迁移后,`RegistryClient` 的实现(~330 行)中会暴露两个问题: + +1. `query_code_symbols` 和 `query_dead_code` 直接在 `impl` 块中拼接 SQL,未复用 `registry::code_symbols` / `registry::dead_code` 子模块。 +2. `RegistryClient` 成为 `registry.rs` 中最臃肿的部分,而该文件原本以数据类型定义为主。 + +**建议的 Phase 2 动作**: + +| 方法 | 当前实现方式 | Phase 2 下沉目标 | +|------|--------------|------------------| +| `query_code_symbols` | 内联 SQL (`SELECT ... FROM code_symbols`) | 新增 `registry::code_symbols::query_code_symbols(conn, ...)` 函数 | +| `query_dead_code` | 内联 SQL (`SELECT ... FROM code_symbols cs ...`) | 新增 `registry::dead_code::query_dead_code(conn, ...)` 函数 | +| `query_call_graph` | 调用 `registry::call_graph::query_call_edges` | ✅ 已下沉,无需改动 | +| `query_dependencies` | 调用 `dependency_graph::*` | ✅ 已下沉,无需改动 | + +下沉后,`impl RegistryClient for AppContext` 将退化为**纯代理层**(类似 `ScanClient`),每个方法仅做: +1. `let conn = self.conn()?;` +2. 调用对应子模块函数; +3. 包装 `serde_json::Value` 返回。 + +这进一步使 `RegistryClient` 实现具备可测试性:业务逻辑可在 `registry` 子模块中直接以 `rusqlite::Connection` 为参数进行单元测试,无需构造完整 `AppContext`。 + +### 2.4 模块结构图(迁移后) + +``` +src/ +├── clients.rs # trait 定义(不变) +├── storage.rs # AppContext 本体 + StorageBackend(~375 行) +├── scan.rs # + impl ScanClient for AppContext +├── health.rs # + impl HealthClient for AppContext +├── sync.rs # + impl SyncClient for AppContext +├── digest.rs # + impl DigestClient for AppContext +├── knowledge_engine/ +│ └── mod.rs # + impl KnowledgeClient for AppContext +├── registry.rs # + impl RegistryClient for AppContext(Phase 2 后退化) +│ ├── code_symbols.rs # Phase 2: 新增 query_code_symbols(conn, ...) +│ ├── dead_code.rs # Phase 2: 新增 query_dead_code(conn, ...) +│ └── ... +``` + +--- + +## 三、向后兼容策略 + +### 3.1 API 兼容性 + +- **`clients.rs` trait 签名**:零变更。所有 MCP tool(`mcp/tools/repo.rs`、`knowledge.rs`、`external.rs`、`code_analysis.rs` 等)和 CLI command(`commands/analysis.rs` 等)的调用代码**无需任何修改**。 +- **`AppContext` 公共字段与方法**:`storage`、`config`、`i18n`、`conn()`、`pool()`、`env_cache()`、`set_env_cache()` 保持公开且语义不变。 +- **`AppContext::with_defaults` / `with_storage`**:构造函数逻辑不变,启动一致性检查(`repair_tantivy_consistency`)保留。 + +### 3.2 唯一破坏性变更(建议纳入 v0.15.0) + +- **`AppContext::conn_mut()`**:与 `conn()` 完全等价(`&mut self` 并未改变 `pool.get()` 的不可变借用语义),存在误导性。建议 **移除** 或标记为 `#[deprecated]`。 + - 影响面:全局搜索显示仅 `commands/` 和 `mcp/tools/` 中极少数代码可能使用。经 `grep` 验证,当前调用方均使用 `ctx.conn()` 或直接 `ctx.pool()`,因此实际影响接近零。 + +### 3.3 编译隔离 + +- 迁移后,`scan.rs` 修改不再导致依赖 `storage.rs` 的模块重新编译(因为 `ScanClient` 实现已离开 `storage.rs`)。 +- `registry.rs` 中的 `impl RegistryClient` 修改不再影响 `StorageBackend` 的编译单元。 + +--- + +## 四、估计工作量和风险 + +### 4.1 涉及文件数量 + +| 阶段 | 变更文件数 | 新增文件数 | 说明 | +|------|------------|------------|------| +| Phase 1 | 7 | 0 | `storage.rs` + 6 个目标模块文件剪切粘贴 | +| Phase 2 | 3 | 0 | `registry.rs`、`registry/code_symbols.rs`、`registry/dead_code.rs` 提取函数 | + +### 4.2 工作量估计 + +- **Phase 1**:**0.5–1 人天**。纯代码迁移,无逻辑变更。主要工作为调整各目标文件顶部的 `use` 语句(引入 `AppContext`、`clients::XxxClient`、`serde_json` 等)。 +- **Phase 2**:**1–1.5 人天**。需将 SQL 逻辑封装为带参数的纯函数,并补充单元测试(使用 `WorkspaceRegistry::init_in_memory()` 构造内存数据库)。 +- **回归测试**:`cargo test` 全量通过 + `cargo clippy` 零警告。由于零行为变更,现有测试即回归测试。 + +### 4.3 风险矩阵 + +| 风险项 | 概率 | 影响 | 缓解措施 | +|--------|------|------|----------| +| `use` 语句循环依赖 | 低 | 中 | 迁移前检查各模块已有的 `use crate::storage` 引用;`AppContext` 本体不依赖任何 Client trait,天然避免循环。 | +| `RegistryClient` 内联 SQL 提取时引入行为偏差 | 中 | 低 | 提取前后对比 SQL 字符串完全一致;对 `query_code_symbols` / `query_dead_code` 补充 in-memory SQLite 单元测试。 | +| 私有字段访问权限 | 低 | 低 | `impl` 块中的代码仅使用 `self.conn()`、`self.pool()`、`self.config`、`self.i18n`、`self.env_cache()`,均为公共 API,跨模块访问无障碍。 | +| 编译时间回退 | 极低 | 低 | 实际上编译时间应略微下降(`storage.rs` 编译单元变小,增量编译更细)。 | + +--- + +## 五、Hard Veto 兼容性检查 + +| Veto 项 | 检查结果 | 说明 | +|---------|----------|------| +| 禁止闭源 / 云端强制 / 数据外泄 | ✅ 兼容 | 纯内部文件移动,无外部依赖引入。 | +| 禁止 Docker / RAG(Qdrant) / GUI(Electron) | ✅ 兼容 | 不涉及。 | +| 禁止项目广度 > 5 核心工具 | ✅ 兼容 | 不新增核心工具或 crate。 | +| 本地 LLM 优先 | ✅ 兼容 | 不涉及模型变更。 | +| Rust 核心模块不可外包给子 Agent | ⚠️ 注意 | 本方案为**设计文档与纯代码迁移**,实际执行建议由人类或 `coder` 子代理在本地完成;若使用子代理,应在人类复核后执行 `git diff`。 | +| 禁止永久删除文件 | ✅ 兼容 | 仅剪切代码,原 `storage.rs` 保留大量代码,不产生废弃文件。 | + +--- + +## 六、执行建议 + +1. **立即执行 Phase 1**:无风险、高回报,直接解决 `storage.rs` 臃肿问题。 +2. **Phase 2 与 v0.15.0 其他 P2 工作并行**:在需要修改 `query_code_symbols` 或 `query_dead_code` 逻辑时顺手提取,不单独立项。 +3. **删除 `conn_mut()`**:在 Phase 1 中一并删除,因为 v0.15.0 已允许 minor breaking changes(且实际调用方为零)。 diff --git a/plans/rf6-unwrap-audit-plan.md b/plans/rf6-unwrap-audit-plan.md new file mode 100644 index 0000000..2e73eef --- /dev/null +++ b/plans/rf6-unwrap-audit-plan.md @@ -0,0 +1,129 @@ +# RF-6 Implementation Plan: Eliminate `unwrap()` / `expect()` / `panic!()` from Production Code + +> **Project:** devbase +> **Rule:** Production code (`src/**/*.rs` outside `#[cfg(test)]` blocks) must have zero `unwrap()`, `expect()`, `panic!()`. +> **Date:** 2026-05-08 +> **Estimated Total Effort:** medium (2–3 hours) + +--- + +## Executive Summary + +Scan found **24 occurrences** across **7 files**. + +**No function signatures need to change** for 21 of the 24 occurrences. +**2 functions require signature changes** in `semantic_index/mod.rs`: +- `index_repo_full` → `anyhow::Result<(Vec, Vec)>` +- `index_repo` → `anyhow::Result>` + +Only **2 callers** need updates for the signature changes. + +--- + +## Execution Order (Lowest Risk First) + +| Order | File | Occurrences | Risk | +|-------|------|-------------|------| +| 1 | `src/test_utils.rs` | 1 | Low | +| 2 | `src/query.rs` | 1 | Low | +| 3 | `src/search.rs` | 12 | Low | +| 4 | `src/search/hybrid.rs` | 2 | Low | +| 5 | `src/workflow/scheduler.rs` | 4 | Low | +| 6 | `src/discovery_engine.rs` | 2 | Low | +| 7 | `src/semantic_index/mod.rs` | 2 | **High** | + +--- + +## Detailed Plan + +### 1. `src/test_utils.rs:9` + +```rust +// BEFORE +pub fn temp_db() -> rusqlite::Connection { + WorkspaceRegistry::init_in_memory().expect("failed to create in-memory db") +} +// AFTER +pub fn temp_db() -> anyhow::Result { + WorkspaceRegistry::init_in_memory() +} +``` + +### 2. `src/discovery_engine.rs:181–182` + +```rust +// BEFORE +let set_a = keywords_map.get(a).expect("repo id from keywords_map keys"); +let set_b = keywords_map.get(b).expect("repo id from keywords_map keys"); +// AFTER +let set_a = keywords_map.get(a).ok_or_else(|| anyhow::anyhow!("repo id {} missing from keywords_map", a))?; +let set_b = keywords_map.get(b).ok_or_else(|| anyhow::anyhow!("repo id {} missing from keywords_map", b))?; +``` + +### 3. `src/query.rs:25` + +```rust +// BEFORE +let first = value.chars().next().expect("value not empty: checked above"); +// AFTER +let first = value.chars().next()?; +``` + +### 4. `src/search.rs` (12 occurrences) + +All `schema.get_field("...").expect("...")` → `schema.get_field("...")?` + +Locations: 97, 99, 103, 105, 109, 128, 164, 267, 271, 272, 276, 298 + +### 5. `src/search/hybrid.rs` + +```rust +// BEFORE +return lists.into_iter().next().expect("lists len == 1 checked above"); +// AFTER +return lists.remove(0); +``` + +```rust +// BEFORE +1 => Ok(lists.into_iter().next().expect("lists len == 1 checked above").into_iter().take(limit).collect()), +// AFTER +1 => Ok(lists.remove(0).into_iter().take(limit).collect()), +``` + +### 6. `src/semantic_index/mod.rs` + +Signature changes: +```rust +pub fn index_repo_full(repo_path: &Path) -> anyhow::Result<(Vec, Vec)> +pub fn index_repo(repo_path: &Path) -> anyhow::Result> { + Ok(index_repo_full(repo_path)?.0) +} +``` + +Line 191: `.expect("failed to spawn index worker")` → `?` +Line 198: `handle.join().unwrap()` → `handle.join().map_err(|e| anyhow::anyhow!("index worker panicked: {:?}", e))?` + +Caller updates: +- `knowledge_engine/index.rs:186` → add `?` +- `semantic_index/mod.rs:274` → wrap with `Ok(...?)` +- Test at `semantic_index/mod.rs:602` → add `.unwrap()` (test exempt) + +### 7. `src/workflow/scheduler.rs` + +| Line | Before | After | +|------|--------|-------| +| 19 | `*in_degree.get_mut(...).expect(...) += 1;` | `let deg = in_degree.get_mut(...).ok_or_else(...)?; *deg += 1;` | +| 36 | `queue.pop_front().expect(...)` | `queue.pop_front().ok_or_else(...)?` | +| 37 | `wf.steps.iter().find(...).expect(...).clone()` | `wf.steps.iter().find(...).ok_or_else(...)?.clone()` | +| 43 | `in_degree.get_mut(child).expect(...)` | `in_degree.get_mut(child).ok_or_else(...)?` | + +--- + +## Verification + +```bash +cargo check +cargo test +cargo clippy -- -D clippy::unwrap_used -D clippy::expect_used +``` diff --git a/plans/v0.15.0-directions.md b/plans/v0.15.0-directions.md new file mode 100644 index 0000000..f8cf6af --- /dev/null +++ b/plans/v0.15.0-directions.md @@ -0,0 +1,91 @@ +# devbase v0.15.0 发展方向调研 + +> 来源:health / index / storage / search / hybrid / embedding / AGENTS.md / invariants.md 代码级审阅 +> 日期:2026-05-08 +> 版本:v0.14.3 → v0.15.0 规划 + +--- + +## 方向 1:Tantivy BM25 代码符号搜索接入(Hybrid Search 升级) + +- **一句话目标**:将 `hybrid.rs` 中 SQLite `LIKE` 降级关键词路径替换为 Tantivy BM25,实现代码符号级全文检索。 +- **当前痛点/机会**: + - `keyword_search_symbols` 仅搜索 `symbol_type = 'function'`,且使用 `LIKE '%token%'` 无法利用索引,大仓库查询延迟高。 + - `search.rs` 的 Tantivy schema 仅索引 repo/vault 级文档(title/content/tags),未下沉到代码符号粒度。 + - AGENTS.md 明确标注 L2 全文搜索"基础设施就绪,待接入 `semantic_search_symbols`"。 +- **涉及主要文件**:`src/search/hybrid.rs`、`src/search.rs`、`src/semantic_index/`、`src/registry/WorkspaceRegistry`(`semantic_search_symbols`) +- **估计工作量**:M +- **Hard Veto 兼容性**:✅ 完全兼容。纯本地 Tantivy,无云端依赖,不引入强制联网。 + +--- + +## 方向 2:Health 模块扩展 — 工具链生态检测与非 Git 工作区细粒度健康 + +- **一句话目标**:扩展 `EnvVersionCache` 覆盖更多本地工具链(Python、Docker、Zig、Bun 等),并为非 Git 工作区提供文件级变更明细而非仅全局 hash。 +- **当前痛点/机会**: + - `refresh_env_cache` 仅检测 rustc/cargo/node/go/cmake 五种工具,现代本地开发环境还依赖 Python、Docker、Bun、Zig、Java 等。 + - 非 Git 工作区(`workspace_type != "git"`)的健康状态只有二元 `changed`/`ok`,用户无法获知具体哪些文件变更。 + - `compute_workspace_hash` 对大目录/二进制资产全量 blake3 哈希,没有大小阈值或采样机制,可能拖慢 health 检查。 +- **涉及主要文件**:`src/health.rs`、`src/registry/workspace.rs`、`src/storage.rs`(`EnvVersionCache`) +- **估计工作量**:S–M +- **Hard Veto 兼容性**:✅ 完全兼容。纯本地子进程探测与文件系统遍历,无数据外泄。 + +--- + +## 方向 3:AppContext 职责拆分与启动自愈强化 + +- **一句话目标**:将 `AppContext` 中 7 个 Client trait 的庞大实现拆分到独立模块,并实现 startup 对 `missing_from_index` 的自动修复(而非仅 warn)。 +- **当前痛点/机会**: + - `storage.rs` 已达 846 行,`AppContext` 同时实现 `ScanClient`、`HealthClient`、`SyncClient`、`DigestClient`、`KnowledgeClient`、`RegistryClient`,违反 SRP 且阻碍单元测试。 + - `repair_tantivy_consistency` 已能检测"SQLite 有但 Tantivy 无"的缺失文档(`missing_from_index`),但仅打印 warn 日志,未触发自动 re-index,用户需手动发现。 + - `conn_mut()` 与 `conn()` 完全等价,存在误导性 API。 +- **涉及主要文件**:`src/storage.rs`、`src/clients.rs`、`src/commands/`、`src/search.rs` +- **估计工作量**:M–L +- **Hard Veto 兼容性**:✅ 完全兼容。纯内部重构,不引入外部依赖。 + +--- + +## 方向 4:Embedding Provider 多后端与离线化 + +- **一句话目标**:在 `devbase-embedding` crate 中增加 Ollama/ONNX 可选 provider,并消除 CandleProvider 首次运行时强制联网下载模型的痛点。 +- **当前痛点/机会**: + - 当前仅 `CandleProvider`(all-MiniLM-L6-v2,384-dim),模型固定不可配置;首次运行通过 `hf_hub::api::sync::Api` 联网下载,无网环境直接失败。 + - AGENTS.md 规划 L1 向量语义使用 768-dim Ollama provider,状态为"待激活"。 + - `EmbeddingProvider` trait 已预留 `encode_batch`,但主代码因 Candle CPU BERT batch 性能问题回滚至 `rayon::par_iter`,需要多后端才能发挥 batch 优势(ONNX/GPU)。 +- **涉及主要文件**:`crates/devbase-embedding/src/lib.rs`、`src/config.rs`(`[embedding]` 配置段)、`src/knowledge_engine/index.rs` +- **估计工作量**:M +- **Hard Veto 兼容性**:⚠️ 需注意。Ollama 为本地可选服务,符合"本地优先";必须避免将云端 API(OpenAI 等)设为默认或强制路径。模型下载应支持离线缓存/预置提示。 + +--- + +## 方向 5:架构不变量自动化 Enforce(CI Fitness Functions) + +- **一句话目标**:将 `invariants.md` 中的 12 条规则(尤其 G5/T11/T12)转化为 CI/本地可自动运行的检测脚本,防止人工审查遗漏。 +- **当前痛点/机会**: + - 不变量清单当前依赖"代码审查"和"人工审查"作为检测方式,无自动化守卫。 + - RF-6(生产代码 unwrap 清零)虽已完成,但无 CI 级自动检查,后续 PR 容易回退。 + - T11(`mcp/tools` 不得直接调用 `rusqlite::Connection`)和 T12(`tui/render` 纯消费者)可通过 `grep`/`ripgrep` 脚本在 CI 中快速 enforce。 + - 模块提取演习(模块拆分健康度)目前是手动表格,可脚本化。 +- **涉及主要文件**:`.github/workflows/ci.yml`、`docs/architecture/invariants.md`、`tools/invariant-checks/`(新增) +- **估计工作量**:S +- **Hard Veto 兼容性**:✅ 完全兼容。纯本地 CI 脚本,无外部服务依赖。 + +--- + +## 优先级建议 + +| 优先级 | 方向 | 理由 | +|--------|------|------| +| P1 | 方向 1(BM25 接入) | AGENTS.md 已标记"基础设施就绪",低垂果实,直接提升搜索性能 | +| P2 | 方向 3(AppContext 拆分) | 技术债中唯一 🔴 级耦合问题,越早拆成本越低 | +| P3 | 方向 4(Embedding 多后端) | 打通 768-dim Ollama 路径,为后续 NLQ 质量提升铺路 | +| P4 | 方向 2(Health 扩展) | 改善日常 CLI 使用体验,工作量可控 | +| P5 | 方向 5(不变量 CI) | 质量守卫基础设施,一次性投入,长期收益 | + +--- + +## 排除项说明(本次不纳入) + +- **tree-sitter 编译成本优化**:属于构建系统债,对运行时功能无直接增益,defer 至独立构建工程 wave。 +- **知识蒸馏 Pipeline**:AGENTS.md 标注为"设计规格级,待验证",过早落地风险高。 +- **SQLite FTS5 替代 Tantivy**:属于 Tantivy+SQLite 双写一致性的长期根治方案,评估工作量大,建议作为 v0.16 独立 ADR。 diff --git a/src/commands/limit.rs b/src/commands/limit.rs index 3195f6e..51464b7 100644 --- a/src/commands/limit.rs +++ b/src/commands/limit.rs @@ -134,6 +134,9 @@ mod tests { fn index_path(&self) -> anyhow::Result { Ok(self.dir.path().join("idx")) } + fn symbol_index_path(&self) -> anyhow::Result { + Ok(self.dir.path().join("sym_idx")) + } fn backup_dir(&self) -> anyhow::Result { Ok(self.dir.path().join("bk")) } diff --git a/src/commands/repo.rs b/src/commands/repo.rs index 6dab778..a013458 100644 --- a/src/commands/repo.rs +++ b/src/commands/repo.rs @@ -7,10 +7,17 @@ pub async fn run_scan( ctx: &mut crate::storage::AppContext, path: &str, register: bool, + json: bool, ) -> anyhow::Result<()> { info!("{}: {}", ctx.i18n.cli.scanning, path); let pool = ctx.pool(); - scan::run(path, register, &pool).await + if json { + let result = scan::run_json(path, register, &pool).await?; + println!("{}", serde_json::to_string_pretty(&result)?); + } else { + scan::run(path, register, &pool).await?; + } + Ok(()) } pub async fn run_health( @@ -21,14 +28,41 @@ pub async fn run_health( json: bool, ) -> anyhow::Result<()> { info!("{}", ctx.i18n.cli.health_check); + + // Refresh env cache if stale + let cache = ctx.env_cache()?; + let env_cache = if cache.is_fresh() { + cache + } else { + let fresh = health::refresh_env_cache().await; + ctx.set_env_cache(fresh.clone())?; + fresh + }; + let conn = ctx.conn()?; if json { - let output = - health::run_json(&conn, detail, limit, page, ctx.config.cache.ttl_seconds, &ctx.i18n) - .await?; + let output = health::run_json( + &conn, + detail, + limit, + page, + ctx.config.cache.ttl_seconds, + &ctx.i18n, + &env_cache, + ) + .await?; println!("{}", serde_json::to_string_pretty(&output)?); } else { - health::run(&conn, detail, limit, page, ctx.config.cache.ttl_seconds, &ctx.i18n).await?; + health::run( + &conn, + detail, + limit, + page, + ctx.config.cache.ttl_seconds, + &ctx.i18n, + &env_cache, + ) + .await?; } Ok(()) } @@ -51,13 +85,17 @@ pub async fn run_query( Ok(()) } -pub async fn run_index(ctx: &mut crate::storage::AppContext, path: &str) -> anyhow::Result<()> { - info!("{}: path='{}'", ctx.i18n.cli.indexing, path); +pub async fn run_index( + ctx: &mut crate::storage::AppContext, + path: &str, + skip_embeddings: bool, +) -> anyhow::Result<()> { + info!("{}: path='{}' skip_embeddings={}", ctx.i18n.cli.indexing, path, skip_embeddings); let path = path.to_string(); let pool = ctx.pool(); let count = tokio::task::spawn_blocking(move || { let mut conn = pool.get()?; - knowledge_engine::run_index(&mut conn, &path) + knowledge_engine::run_index(&mut conn, &path, skip_embeddings) }) .await .map_err(|e| anyhow::anyhow!("spawn_blocking failed: {}", e))??; @@ -367,6 +405,60 @@ pub async fn run_status(ctx: &mut crate::storage::AppContext, json: bool) -> any Ok(()) } +pub fn run_knowledge_report( + ctx: &mut crate::storage::AppContext, + repo_id: &str, + activity_limit: usize, + json: bool, +) -> anyhow::Result<()> { + let conn = ctx.conn()?; + let report = crate::oplog_analytics::generate_report( + &conn, + if repo_id.is_empty() { + None + } else { + Some(repo_id) + }, + activity_limit, + )?; + if json { + println!("{}", serde_json::to_string_pretty(&report)?); + } else { + println!("Knowledge Coverage Report"); + println!("========================="); + println!( + "Repos: {} | Symbols: {} | Embeddings: {} | Calls: {}", + report.repo_count, report.total_symbols, report.total_embeddings, report.total_calls + ); + println!("Overall coverage: {:.1}%", report.overall_coverage_pct); + if !report.repos.is_empty() { + println!("\nPer-repo breakdown:"); + for r in &report.repos { + println!( + " [{}] symbols={} embeddings={} calls={} coverage={:.1}%", + r.repo_id, r.symbol_count, r.embedding_count, r.call_count, r.coverage_pct + ); + } + } + println!( + "\nHealth summary: dirty={} ahead={} behind={} diverged={} up_to_date={}", + report.health_summary.dirty, + report.health_summary.ahead, + report.health_summary.behind, + report.health_summary.diverged, + report.health_summary.up_to_date + ); + if !report.recent_activity.is_empty() { + println!("\nRecent activity (last {} events):", report.recent_activity.len()); + for act in &report.recent_activity { + let repo = act.repo_id.as_deref().unwrap_or("-"); + println!(" [{}] repo={} type={}", act.timestamp, repo, act.event_type); + } + } + } + Ok(()) +} + pub fn run_registry( ctx: &mut crate::storage::AppContext, cmd: crate::RegistryCommands, diff --git a/src/commands/simple.rs b/src/commands/simple.rs index 5bc5ec3..ff466f4 100644 --- a/src/commands/simple.rs +++ b/src/commands/simple.rs @@ -13,8 +13,8 @@ pub use crate::commands::knowledge::{ #[cfg(feature = "tui")] pub use crate::commands::repo::run_discover; pub use crate::commands::repo::{ - run_health, run_index, run_query, run_registry, run_scan, run_status, run_sync, - run_syncthing_push, + run_health, run_index, run_knowledge_report, run_query, run_registry, run_scan, run_status, + run_sync, run_syncthing_push, }; #[cfg(feature = "mcp")] pub use crate::commands::system::run_mcp; @@ -53,6 +53,9 @@ mod tests { fn index_path(&self) -> anyhow::Result { Ok(self.dir.path().join("idx")) } + fn symbol_index_path(&self) -> anyhow::Result { + Ok(self.dir.path().join("sym_idx")) + } fn backup_dir(&self) -> anyhow::Result { Ok(self.dir.path().join("bk")) } diff --git a/src/commands/skill.rs b/src/commands/skill.rs index d0d4f16..bb0b1b2 100644 --- a/src/commands/skill.rs +++ b/src/commands/skill.rs @@ -434,6 +434,9 @@ mod tests { fn index_path(&self) -> anyhow::Result { Ok(self.dir.path().join("idx")) } + fn symbol_index_path(&self) -> anyhow::Result { + Ok(self.dir.path().join("sym_idx")) + } fn backup_dir(&self) -> anyhow::Result { Ok(self.dir.path().join("bk")) } diff --git a/src/commands/workflow.rs b/src/commands/workflow.rs index 0d987db..ee2a8c3 100644 --- a/src/commands/workflow.rs +++ b/src/commands/workflow.rs @@ -6,14 +6,35 @@ pub fn run_workflow( ) -> anyhow::Result<()> { let conn = ctx.conn_mut()?; match cmd { - crate::WorkflowCommands::List => { + crate::WorkflowCommands::List { json } => { let workflows = crate::workflow::list_workflows(&conn)?; - if workflows.is_empty() { - println!("No workflows registered."); + if json { + let items: Vec = workflows + .into_iter() + .map(|(id, name, version)| { + serde_json::json!({ + "id": id, + "name": name, + "version": version + }) + }) + .collect(); + println!( + "{}", + serde_json::to_string_pretty(&serde_json::json!({ + "success": true, + "count": items.len(), + "workflows": items + }))? + ); } else { - println!("Registered workflows:"); - for (id, name, version) in workflows { - println!(" [{}] {} (v{})", id, name, version); + if workflows.is_empty() { + println!("No workflows registered."); + } else { + println!("Registered workflows:"); + for (id, name, version) in workflows { + println!(" [{}] {} (v{})", id, name, version); + } } } } @@ -142,6 +163,9 @@ mod tests { fn index_path(&self) -> anyhow::Result { Ok(self.dir.path().join("idx")) } + fn symbol_index_path(&self) -> anyhow::Result { + Ok(self.dir.path().join("sym_idx")) + } fn backup_dir(&self) -> anyhow::Result { Ok(self.dir.path().join("bk")) } @@ -151,7 +175,7 @@ mod tests { fn test_run_workflow_list_empty() { let storage = Arc::new(TempStorage::new()); let mut ctx = AppContext::with_storage(storage).unwrap(); - let result = run_workflow(&mut ctx, crate::WorkflowCommands::List); + let result = run_workflow(&mut ctx, crate::WorkflowCommands::List { json: false }); assert!(result.is_ok()); } diff --git a/src/config.rs b/src/config.rs index 99adada..fcedb99 100644 --- a/src/config.rs +++ b/src/config.rs @@ -250,7 +250,7 @@ fn default_embedding_provider() -> String { "ollama".to_string() } fn default_embedding_model() -> String { - "nomic-embed-text".to_string() + "all-minilm".to_string() } fn default_embedding_base_url() -> String { "http://localhost:11434".to_string() @@ -376,10 +376,12 @@ max_tokens = 200 timeout_seconds = 30 [embedding] -# Local embedding for semantic code search. Requires Ollama installed. +# Local embedding for semantic code search. +# Backend: "candle" (pure Rust, all-MiniLM-L6-v2) or "ollama" (requires local Ollama). +# Use "all-minilm" model with Ollama for 384-dim embeddings (compatible with candle). enabled = false provider = "ollama" -model = "nomic-embed-text" +model = "all-minilm" base_url = "http://localhost:11434" timeout_seconds = 30 diff --git a/src/digest.rs b/src/digest.rs index 9a90c10..d57a256 100644 --- a/src/digest.rs +++ b/src/digest.rs @@ -117,6 +117,14 @@ pub fn generate_daily_digest( Ok(lines.join("\n")) } +impl crate::clients::DigestClient for crate::storage::AppContext { + fn generate_daily_digest(&self) -> anyhow::Result { + let conn = self.conn()?; + let text = crate::digest::generate_daily_digest(&conn, &self.config, &self.i18n)?; + Ok(serde_json::json!({ "success": true, "digest": text })) + } +} + #[cfg(test)] mod tests { use super::*; diff --git a/src/discovery_engine.rs b/src/discovery_engine.rs index 8f3aa29..a148a58 100644 --- a/src/discovery_engine.rs +++ b/src/discovery_engine.rs @@ -177,9 +177,12 @@ pub fn discover_similar_projects(conn: &rusqlite::Connection) -> anyhow::Result< for j in (i + 1)..repo_ids.len() { let a = &repo_ids[i]; let b = &repo_ids[j]; - // TODO(veto-audit-2026-04-26): RF-6 expect — Map 内部状态可能不一致,应返回 Option 或默认值。 - let set_a = keywords_map.get(a).expect("repo id from keywords_map keys"); - let set_b = keywords_map.get(b).expect("repo id from keywords_map keys"); + let set_a = keywords_map + .get(a) + .ok_or_else(|| anyhow::anyhow!("repo id from keywords_map keys"))?; + let set_b = keywords_map + .get(b) + .ok_or_else(|| anyhow::anyhow!("repo id from keywords_map keys"))?; let intersection: HashSet = set_a.intersection(set_b).cloned().collect(); if intersection.is_empty() { diff --git a/src/embedding.rs b/src/embedding.rs index 5a6dd22..cc0d6d3 100644 --- a/src/embedding.rs +++ b/src/embedding.rs @@ -6,6 +6,30 @@ #[cfg(feature = "embedding")] pub use devbase_embedding::*; +#[cfg(feature = "embedding")] +static CONFIG_PROVIDER: std::sync::OnceLock> = + std::sync::OnceLock::new(); + +/// Generate a query embedding, respecting the user's embedding backend configuration. +/// Falls back to the default Candle provider if config cannot be loaded. +#[cfg(feature = "embedding")] +pub fn generate_query_embedding(text: &str) -> anyhow::Result> { + let provider = CONFIG_PROVIDER.get_or_init(|| { + crate::config::Config::load() + .ok() + .map(|c| { + create_provider( + &c.embedding.provider, + &c.embedding.model, + &c.embedding.base_url, + c.embedding.timeout_seconds, + ) + }) + .unwrap_or_else(default_provider) + }); + provider.encode(text) +} + #[cfg(not(feature = "embedding"))] pub fn generate_query_embedding(_text: &str) -> anyhow::Result> { anyhow::bail!("Embedding support is disabled. Enable the 'embedding' feature.") diff --git a/src/health.rs b/src/health.rs index 78870f8..6234b73 100644 --- a/src/health.rs +++ b/src/health.rs @@ -60,6 +60,7 @@ pub async fn run_json( page: usize, ttl_seconds: i64, i18n: &crate::i18n::I18n, + env_cache: &crate::storage::EnvVersionCache, ) -> anyhow::Result { let start = std::time::Instant::now(); let (total_repos, dirty_repos, behind_upstream, no_upstream_count, repo_details) = { @@ -85,6 +86,15 @@ pub async fn run_json( let mut no_upstream_count: usize = 0; let mut repo_details: Vec = Vec::new(); + // Batch load health cache for all git repos + let repo_ids: Vec<&str> = repos + .iter() + .filter(|r| r.workspace_type == "git") + .map(|r| r.id.as_str()) + .collect(); + let health_batch = + crate::registry::health::get_health_batch(conn, &repo_ids).unwrap_or_default(); + for repo in repos { total_repos += 1; let primary = repo.primary_remote(); @@ -92,11 +102,12 @@ pub async fn run_json( let default_branch = primary.and_then(|r| r.default_branch.clone()); let (status, ahead, behind) = if repo.workspace_type == "git" { - match crate::registry::health::get_health(conn, &repo.id) { - Ok(Some(health)) => { + match health_batch.get(&repo.id).cloned() { + Some(health) => { let elapsed = Utc::now().signed_duration_since(health.checked_at).num_seconds(); - if elapsed < ttl_seconds { + // Guard against negative elapsed (system clock drift / future checked_at) + if elapsed >= 0 && elapsed < ttl_seconds { (health.status, health.ahead, health.behind) } else { let (status, ahead, behind) = analyze_repo( @@ -118,7 +129,7 @@ pub async fn run_json( (status, ahead, behind) } } - _ => { + None => { let (status, ahead, behind) = analyze_repo( repo.local_path.to_string_lossy().as_ref(), upstream_url.as_deref(), @@ -206,11 +217,15 @@ pub async fn run_json( }; let environment = serde_json::json!({ - "rustc": get_tool_version("rustc", &["--version"]).await.map(|s| fmt_version(Some(s), i18n)), - "cargo": get_tool_version("cargo", &["--version"]).await.map(|s| fmt_version(Some(s), i18n)), - "node": get_tool_version("node", &["--version"]).await.map(|s| fmt_version(Some(s), i18n)), - "go": get_tool_version("go", &["version"]).await.map(|s| fmt_version(Some(s), i18n)), - "cmake": get_tool_version("cmake", &["--version"]).await.map(|s| fmt_version(Some(s), i18n)), + "rustc": fmt_version(env_cache.rustc.clone(), i18n), + "cargo": fmt_version(env_cache.cargo.clone(), i18n), + "node": fmt_version(env_cache.node.clone(), i18n), + "go": fmt_version(env_cache.go.clone(), i18n), + "cmake": fmt_version(env_cache.cmake.clone(), i18n), + "python": fmt_version(env_cache.python.clone(), i18n), + "bun": fmt_version(env_cache.bun.clone(), i18n), + "zig": fmt_version(env_cache.zig.clone(), i18n), + "java": fmt_version(env_cache.java.clone(), i18n), }); let summary = serde_json::json!({ @@ -277,8 +292,9 @@ pub async fn run( page: usize, ttl_seconds: i64, i18n: &crate::i18n::I18n, + env_cache: &crate::storage::EnvVersionCache, ) -> anyhow::Result<()> { - let result = run_json(conn, detail, limit, page, ttl_seconds, i18n).await?; + let result = run_json(conn, detail, limit, page, ttl_seconds, i18n, env_cache).await?; let summary = result["summary"] .as_object() @@ -298,6 +314,10 @@ pub async fn run( println!(" node: {}", env["node"].as_str().unwrap_or(i18n.log.not_installed)); println!(" go: {}", env["go"].as_str().unwrap_or(i18n.log.not_installed)); println!(" cmake: {}", env["cmake"].as_str().unwrap_or(i18n.log.not_installed)); + println!(" python: {}", env["python"].as_str().unwrap_or(i18n.log.not_installed)); + println!(" bun: {}", env["bun"].as_str().unwrap_or(i18n.log.not_installed)); + println!(" zig: {}", env["zig"].as_str().unwrap_or(i18n.log.not_installed)); + println!(" java: {}", env["java"].as_str().unwrap_or(i18n.log.not_installed)); if detail { let repos = result["repos"] @@ -372,16 +392,16 @@ pub fn analyze_repo( } // Check for detached HEAD - let is_detached = match repo.head() { - Ok(head) => head.target().is_none(), - Err(_) => true, + let head = match repo.head() { + Ok(h) => h, + Err(_) => return ("detached".to_string(), 0, 0), }; - if is_detached { + if head.target().is_none() { return ("detached".to_string(), 0, 0); } - let (ahead, behind) = match calc_ahead_behind(&repo, default_branch) { + let (ahead, behind) = match calc_ahead_behind(&repo, head, default_branch) { Ok(ab) => ab, Err(_) => return ("ok".to_string(), 0, 0), }; @@ -403,13 +423,9 @@ pub fn analyze_repo( fn calc_ahead_behind( repo: &Repository, + head: git2::Reference, default_branch: Option<&str>, ) -> anyhow::Result<(usize, usize)> { - let head = match repo.head() { - Ok(h) => h, - Err(_) => return Ok((0, 0)), - }; - let local_oid = match head.target() { Some(oid) => oid, None => return Ok((0, 0)), @@ -441,6 +457,33 @@ fn calc_ahead_behind( repo.graph_ahead_behind(local_oid, remote_oid).map_err(|e| anyhow::anyhow!(e)) } +/// Refresh environment version cache by spawning all tool subprocesses in parallel. +pub async fn refresh_env_cache() -> crate::storage::EnvVersionCache { + let (rustc, cargo, node, go, cmake, python, bun, zig, java) = tokio::join!( + get_tool_version("rustc", &["--version"]), + get_tool_version("cargo", &["--version"]), + get_tool_version("node", &["--version"]), + get_tool_version("go", &["version"]), + get_tool_version("cmake", &["--version"]), + get_tool_version("python", &["--version"]), + get_tool_version("bun", &["--version"]), + get_tool_version("zig", &["version"]), + get_tool_version("java", &["-version"]), + ); + crate::storage::EnvVersionCache { + rustc, + cargo, + node, + go, + cmake, + python, + bun, + zig, + java, + fetched_at: Some(std::time::Instant::now()), + } +} + async fn get_tool_version(cmd: &str, args: &[&str]) -> Option { let output = tokio::process::Command::new(cmd).args(args).output().await.ok()?; @@ -448,7 +491,11 @@ async fn get_tool_version(cmd: &str, args: &[&str]) -> Option { return None; } - let raw = String::from_utf8_lossy(&output.stdout); + let raw = if !output.stdout.is_empty() { + String::from_utf8_lossy(&output.stdout) + } else { + String::from_utf8_lossy(&output.stderr) + }; let line = raw.lines().next()?.trim(); if line.is_empty() { return None; @@ -459,27 +506,63 @@ async fn get_tool_version(cmd: &str, args: &[&str]) -> Option { fn fmt_version(raw: Option, i18n: &crate::i18n::I18n) -> String { match raw { Some(s) => { + let s = s.trim(); + // Java outputs: java version "1.8.0_31" — extract quoted version + if let Some(start) = s.find('"') { + if let Some(end) = s[start + 1..].find('"') { + return s[start + 1..start + 1 + end].to_string(); + } + } let parts: Vec<&str> = s.split_whitespace().collect(); if parts.len() >= 2 { match parts[0] { - "rustc" | "cargo" => parts.get(1).unwrap_or(&"unknown").to_string(), + "rustc" | "cargo" | "bun" | "zig" | "Python" => { + parts.get(1).unwrap_or(&"unknown").to_string() + } "cmake" | "version" => parts.get(2).unwrap_or(&"unknown").to_string(), + "go" if parts.len() >= 3 => parts[2].to_string(), + "Docker" if parts.len() >= 3 && parts[1] == "version" => parts[2..].join(" "), _ => { - if parts[0] == "go" && parts.len() >= 3 { - parts[2].to_string() + // Heuristic: if second word is "version", skip first two + if parts.len() >= 3 && parts[1] == "version" { + parts[2..].join(" ") } else { - s + s.to_string() } } } } else { - s + s.to_string() } } None => i18n.log.not_installed.to_string(), } } +impl crate::clients::HealthClient for crate::storage::AppContext { + async fn check_health(&self, detail: bool) -> anyhow::Result { + let conn = self.conn()?; + let cache = self.env_cache()?; + let env_cache = if cache.is_fresh() { + cache + } else { + let fresh = crate::health::refresh_env_cache().await; + self.set_env_cache(fresh.clone())?; + fresh + }; + crate::health::run_json( + &conn, + detail, + 0, + 1, + self.config.cache.ttl_seconds, + &self.i18n, + &env_cache, + ) + .await + } +} + #[cfg(test)] mod tests { use super::*; @@ -566,4 +649,31 @@ mod tests { let i18n = crate::i18n::from_language("en"); assert_eq!(fmt_version(Some("v1.0".to_string()), &i18n), "v1.0"); } + + #[test] + fn test_fmt_version_python() { + let i18n = crate::i18n::from_language("en"); + assert_eq!(fmt_version(Some("Python 3.12.13".to_string()), &i18n), "3.12.13"); + } + + #[test] + fn test_fmt_version_bun() { + let i18n = crate::i18n::from_language("en"); + assert_eq!(fmt_version(Some("1.3.11".to_string()), &i18n), "1.3.11"); + } + + #[test] + fn test_fmt_version_java() { + let i18n = crate::i18n::from_language("en"); + assert_eq!(fmt_version(Some("java version \"1.8.0_31\"".to_string()), &i18n), "1.8.0_31"); + } + + #[test] + fn test_fmt_version_docker() { + let i18n = crate::i18n::from_language("en"); + assert_eq!( + fmt_version(Some("Docker version 24.0.7, build afdd53b".to_string()), &i18n), + "24.0.7, build afdd53b" + ); + } } diff --git a/src/knowledge_engine/index.rs b/src/knowledge_engine/index.rs index d511f70..2c59804 100644 --- a/src/knowledge_engine/index.rs +++ b/src/knowledge_engine/index.rs @@ -1,7 +1,6 @@ // SPDX-License-Identifier: MIT // Copyright (c) 2026 juice094 use crate::registry::RepoEntry; -use rayon::prelude::*; use std::path::PathBuf; fn index_repo_in_search( @@ -59,8 +58,12 @@ pub fn index_repo( } /// 兼容旧调用的包装层:执行索引逻辑 -pub fn run_index(conn: &mut rusqlite::Connection, path: &str) -> anyhow::Result { - run_index_with_progress(conn, path, None) +pub fn run_index( + conn: &mut rusqlite::Connection, + path: &str, + skip_embeddings: bool, +) -> anyhow::Result { + run_index_with_progress(conn, path, None, skip_embeddings) } /// 带进度上报的索引逻辑。 @@ -69,6 +72,7 @@ pub fn run_index_with_progress( conn: &mut rusqlite::Connection, path: &str, progress_tx: Option>, + skip_embeddings: bool, ) -> anyhow::Result { use tracing::{info, warn}; @@ -110,6 +114,7 @@ pub fn run_index_with_progress( let mut count = 0; for repo in &repos { + let t0 = std::time::Instant::now(); let config = crate::config::Config::load().ok(); let (summary, keywords) = config .as_ref() @@ -121,8 +126,10 @@ pub fn run_index_with_progress( warn!("No README found for {}, generating fallback summary", repo.id); super::generate_fallback_summary(&repo.local_path) }); + let t1 = std::time::Instant::now(); let modules = super::extract_module_structure(&repo.local_path); + let t2 = std::time::Instant::now(); crate::registry::knowledge::save_summary(conn, &repo.id, &summary, &keywords)?; @@ -136,6 +143,7 @@ pub fn run_index_with_progress( &keywords, &repo.tags, )?; + let t3 = std::time::Instant::now(); let modules_tuple: Vec<(String, String)> = modules.into_iter().map(|m| (m.name, m.kind)).collect(); @@ -175,8 +183,9 @@ pub fn run_index_with_progress( crate::semantic_index::index_repo_incremental(&repo.local_path, changed) } else { // Full index - crate::semantic_index::index_repo_full(&repo.local_path) + crate::semantic_index::index_repo_full(&repo.local_path)? }; + let t4 = std::time::Instant::now(); if !symbols.is_empty() { let result = if is_incremental { @@ -191,6 +200,13 @@ pub fn run_index_with_progress( } Err(e) => warn!("Failed to save code symbols for {}: {}", repo.id, e), } + + // Index symbols in Tantivy for BM25 keyword search + if let Err(e) = index_symbols_in_search(&repo.id, &symbols, is_incremental) { + warn!("Failed to index symbols in search for {}: {}", repo.id, e); + } else { + notify(format!("symbol_index:{},symbols={}", repo.id, symbols.len())); + } } if !calls.is_empty() { let result = if is_incremental { @@ -206,9 +222,10 @@ pub fn run_index_with_progress( Err(e) => warn!("Failed to save call graph for {}: {}", repo.id, e), } } + let t5 = std::time::Instant::now(); // Generate embeddings for code symbols (local candle, Sprint 14) - if !symbols.is_empty() { + if !skip_embeddings && !symbols.is_empty() { let result = if is_incremental { save_symbol_embeddings_incremental(conn, &repo.id, &symbols) } else { @@ -222,6 +239,7 @@ pub fn run_index_with_progress( Err(e) => warn!("Failed to save symbol embeddings for {}: {}", repo.id, e), } } + let t6 = std::time::Instant::now(); // Save repo_index_state for next incremental run if let Ok(Some(hash)) = crate::semantic_index::git_diff::current_head_hash(&repo.local_path) @@ -239,6 +257,7 @@ pub fn run_index_with_progress( } Err(e) => warn!("Failed to build dependency graph for {}: {}", repo.id, e), } + let t7 = std::time::Instant::now(); println!( "Indexed [{}] -> \"{}\" (keywords: {}) language={:?} symbols={} calls={}", @@ -249,6 +268,17 @@ pub fn run_index_with_progress( symbols.len(), calls.len(), ); + println!( + " timings: readme={:.0}ms module={:.0}ms tantivy={:.0}ms semantic={:.0}ms save={:.0}ms embed={:.0}ms deps={:.0}ms total={:.0}ms", + (t1 - t0).as_millis(), + (t2 - t1).as_millis(), + (t3 - t2).as_millis(), + (t4 - t3).as_millis(), + (t5 - t4).as_millis(), + (t6 - t5).as_millis(), + (t7 - t6).as_millis(), + (t7 - t0).as_millis(), + ); count += 1; } @@ -280,9 +310,13 @@ fn generate_and_save_embeddings( symbols: &[crate::semantic_index::CodeSymbol], clear_existing: bool, ) -> anyhow::Result { + use rayon::prelude::*; use tracing::{info, warn}; - // Phase 1: parallel encoding + // Phase 1: parallel encoding (rayon par_iter gives best throughput for + // Candle CPU BERT because per-symbol sequences are short and variable; + // batching causes excessive padding and Candle's CPU matmul is slower + // for large padded batches than many small single inferences). let items: Vec<(String, String, Vec)> = symbols .par_iter() .filter_map(|sym| { @@ -387,6 +421,19 @@ fn detect_changes( } } +fn index_symbols_in_search( + repo_id: &str, + symbols: &[crate::semantic_index::CodeSymbol], + _is_incremental: bool, +) -> anyhow::Result<()> { + let (index, _reader) = crate::search::symbol_index::init_index()?; + let mut writer = crate::search::symbol_index::get_writer(&index)?; + let schema = index.schema(); + crate::search::symbol_index::add_symbols(&mut writer, &schema, repo_id, symbols)?; + crate::search::symbol_index::commit_writer(&mut writer)?; + Ok(()) +} + fn save_repo_index_state( conn: &mut rusqlite::Connection, repo_id: &str, diff --git a/src/knowledge_engine/mod.rs b/src/knowledge_engine/mod.rs index e81be40..f482e79 100644 --- a/src/knowledge_engine/mod.rs +++ b/src/knowledge_engine/mod.rs @@ -44,3 +44,50 @@ where } } } + +impl crate::clients::KnowledgeClient for crate::storage::AppContext { + fn run_index(&self, path: &str) -> anyhow::Result { + let mut conn = self.conn()?; + let count = crate::knowledge_engine::run_index(&mut conn, path, false)?; + Ok(serde_json::json!({ "success": true, "indexed": count, "errors": 0 })) + } + + fn save_note( + &self, + repo_id: &str, + text: &str, + author: &str, + ) -> anyhow::Result { + let conn = self.conn()?; + crate::registry::knowledge::save_note(&conn, repo_id, text, author)?; + Ok(serde_json::json!({ "success": true })) + } + + fn save_summary( + &self, + repo_id: &str, + desc: &str, + author: &str, + ) -> anyhow::Result { + let conn = self.conn()?; + crate::registry::knowledge::save_summary(&conn, repo_id, desc, author)?; + Ok(serde_json::json!({ "success": true })) + } + + fn get_paper(&self, arxiv_id: &str) -> anyhow::Result { + let conn = self.conn()?; + let papers = crate::registry::knowledge::list_papers(&conn)?; + match papers.into_iter().find(|p| p.id == arxiv_id) { + Some(p) => Ok(serde_json::json!({ + "success": true, + "id": p.id, + "title": p.title, + "venue": p.venue, + "year": p.year, + "pdf_path": p.pdf_path, + "tags": p.tags, + })), + None => Ok(serde_json::json!({ "success": false, "error": "Paper not found" })), + } + } +} diff --git a/src/main.rs b/src/main.rs index 893327f..2de18dc 100644 --- a/src/main.rs +++ b/src/main.rs @@ -24,6 +24,9 @@ pub(crate) enum Commands { /// Register discovered repos into the database #[arg(long)] register: bool, + /// Output results as JSON + #[arg(long)] + json: bool, }, /// Check the health of registered repositories and the environment Health { @@ -80,6 +83,9 @@ pub(crate) enum Commands { /// Specific path to index; if omitted, index all registered repos #[arg(default_value = "")] path: String, + /// Skip semantic embedding generation (symbols/calls still indexed) + #[arg(long)] + skip_embeddings: bool, }, /// Remove archive/backup entries from registry Clean, @@ -160,6 +166,18 @@ pub(crate) enum Commands { Discover, /// Generate daily knowledge digest Digest, + /// Generate knowledge coverage report for the workspace or a repo + KnowledgeReport { + /// Specific repo ID; if omitted, reports on the entire workspace + #[arg(default_value = "")] + repo_id: String, + /// Number of recent activity events to include + #[arg(long, default_value_t = 20)] + activity_limit: usize, + /// Output as JSON + #[arg(long)] + json: bool, + }, /// View the operation log Oplog { /// Limit number of entries (default: 20) @@ -420,7 +438,11 @@ pub(crate) enum SkillCommands { #[derive(Subcommand)] pub(crate) enum WorkflowCommands { /// List registered workflows - List, + List { + /// Output as JSON + #[arg(long)] + json: bool, + }, /// Show workflow definition Show { /// Workflow ID @@ -583,8 +605,8 @@ async fn main() -> anyhow::Result<()> { let cli = Cli::parse(); match cli.command { - Commands::Scan { path, register } => { - commands::simple::run_scan(&mut ctx, &path, register).await?; + Commands::Scan { path, register, json } => { + commands::simple::run_scan(&mut ctx, &path, register, json).await?; } Commands::Health { detail, limit, page, json } => { commands::simple::run_health(&mut ctx, detail, limit, page, json).await?; @@ -603,8 +625,8 @@ async fn main() -> anyhow::Result<()> { Commands::Query { query, limit, page, json } => { commands::simple::run_query(&mut ctx, &query, limit, page, json).await?; } - Commands::Index { path } => { - commands::simple::run_index(&mut ctx, &path).await?; + Commands::Index { path, skip_embeddings } => { + commands::simple::run_index(&mut ctx, &path, skip_embeddings).await?; } Commands::Clean => { commands::simple::run_clean(&mut ctx)?; @@ -731,6 +753,9 @@ async fn main() -> anyhow::Result<()> { Commands::Workflow { cmd } => { commands::workflow::run_workflow(&mut ctx, cmd)?; } + Commands::KnowledgeReport { repo_id, activity_limit, json } => { + commands::simple::run_knowledge_report(&mut ctx, &repo_id, activity_limit, json)?; + } Commands::Limit { cmd } => { commands::limit::run_limit(&mut ctx, cmd)?; } diff --git a/src/mcp/tools/status.rs b/src/mcp/tools/status.rs index 54e6ea5..a32a5db 100644 --- a/src/mcp/tools/status.rs +++ b/src/mcp/tools/status.rs @@ -171,7 +171,7 @@ returns a JSON array of progress events."#, let pool = ctx.pool(); let count = tokio::task::spawn_blocking(move || { let mut conn = pool.get()?; - crate::knowledge_engine::run_index(&mut conn, &path) + crate::knowledge_engine::run_index(&mut conn, &path, false) }) .await .map_err(|e| anyhow::anyhow!("spawn_blocking failed: {}", e))??; @@ -199,7 +199,12 @@ returns a JSON array of progress events."#, let path_clone = path.clone(); let handle = tokio::task::spawn_blocking(move || { let mut conn = pool.get()?; - crate::knowledge_engine::run_index_with_progress(&mut conn, &path_clone, Some(tx)) + crate::knowledge_engine::run_index_with_progress( + &mut conn, + &path_clone, + Some(tx), + false, + ) }); // Collect progress events as they arrive from the blocking thread. diff --git a/src/query.rs b/src/query.rs index 6007870..b8879d4 100644 --- a/src/query.rs +++ b/src/query.rs @@ -21,8 +21,7 @@ pub(crate) fn parse_cmp_expr(value: &str) -> Option<(String, i64)> { if value.is_empty() { return None; } - // TODO(veto-audit-2026-04-26): RF-6 expect — 前置检查存在但非类型级保证,建议改为 `value.chars().next().unwrap_or('\0')` 或 `ok_or`。 - let first = value.chars().next().expect("value not empty: checked above"); + let first = value.chars().next()?; if first == '>' || first == '<' || first == '=' { let num = value[1..].parse().ok()?; Some((first.to_string(), num)) diff --git a/src/registry.rs b/src/registry.rs index 185360c..0d89121 100644 --- a/src/registry.rs +++ b/src/registry.rs @@ -120,6 +120,270 @@ pub mod repos_toml; pub mod vault; pub mod workspace; +impl crate::clients::RegistryClient for crate::storage::AppContext { + fn list_repos(&self, _filter: Option<&str>) -> anyhow::Result { + let conn = self.conn()?; + let repos = crate::registry::repo::list_repos(&conn)?; + let results: Vec = repos + .into_iter() + .map(|r| { + serde_json::json!({ + "id": r.id, + "local_path": r.local_path, + "language": r.language, + "tags": r.tags, + "workspace_type": r.workspace_type, + "data_tier": r.data_tier, + }) + }) + .collect(); + Ok(serde_json::json!({ "success": true, "count": results.len(), "repos": results })) + } + + fn get_repo(&self, repo_id: &str) -> anyhow::Result { + let conn = self.conn()?; + let repos = crate::registry::repo::list_repos(&conn)?; + match repos.into_iter().find(|r| r.id == repo_id) { + Some(r) => Ok(serde_json::json!({ + "success": true, + "id": r.id, + "local_path": r.local_path, + "language": r.language, + "tags": r.tags, + "workspace_type": r.workspace_type, + "data_tier": r.data_tier, + })), + None => Ok(serde_json::json!({ "success": false, "error": "repo not found" })), + } + } + + fn list_modules(&self, repo_id: &str) -> anyhow::Result { + let conn = self.conn()?; + let modules = crate::registry::knowledge::list_modules(&conn, repo_id)?; + let results: Vec = modules + .into_iter() + .map(|(name, ty, path)| { + serde_json::json!({ + "name": name, + "type": ty, + "path": path, + }) + }) + .collect(); + Ok(serde_json::json!({ "success": true, "count": results.len(), "modules": results })) + } + + fn save_paper(&self, paper: &serde_json::Value) -> anyhow::Result { + let conn = self.conn()?; + let paper_entry: crate::registry::PaperEntry = serde_json::from_value(paper.clone())?; + crate::registry::knowledge::save_paper(&conn, &paper_entry)?; + Ok(serde_json::json!({ "success": true })) + } + + fn save_experiment(&self, exp: &serde_json::Value) -> anyhow::Result { + let conn = self.conn()?; + let exp_entry: crate::registry::ExperimentEntry = serde_json::from_value(exp.clone())?; + crate::registry::WorkspaceRegistry::save_experiment(&conn, &exp_entry)?; + Ok(serde_json::json!({ "success": true })) + } + + fn list_code_metrics(&self) -> anyhow::Result { + let conn = self.conn()?; + let metrics = crate::registry::metrics::list_code_metrics(&conn)?; + let repos: Vec = metrics + .into_iter() + .map(|(id, m)| { + serde_json::json!({ + "repo_id": id, + "total_lines": m.total_lines, + "source_lines": m.source_lines, + "test_lines": m.test_lines, + "comment_lines": m.comment_lines, + "file_count": m.file_count, + "language_breakdown": m.language_breakdown, + "updated_at": m.updated_at.to_rfc3339() + }) + }) + .collect(); + Ok(serde_json::json!({ "success": true, "count": repos.len(), "repos": repos })) + } + + fn get_code_metrics(&self, repo_id: &str) -> anyhow::Result { + let conn = self.conn()?; + match crate::registry::metrics::get_code_metrics(&conn, repo_id)? { + Some(m) => Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "total_lines": m.total_lines, + "source_lines": m.source_lines, + "test_lines": m.test_lines, + "comment_lines": m.comment_lines, + "file_count": m.file_count, + "language_breakdown": m.language_breakdown, + "updated_at": m.updated_at.to_rfc3339() + })), + None => { + Ok(serde_json::json!({ "success": false, "error": "No metrics found for repo" })) + } + } + } + + fn get_health(&self, repo_id: &str) -> anyhow::Result { + let conn = self.conn()?; + match crate::registry::health::get_health(&conn, repo_id)? { + Some(h) => Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "status": h.status, + "ahead": h.ahead, + "behind": h.behind, + "checked_at": h.checked_at.to_rfc3339() + })), + None => Ok(serde_json::json!({ "success": false, "error": "No health data found" })), + } + } + + fn query_call_graph( + &self, + repo_id: &str, + callee: Option<&str>, + caller: Option<&str>, + file: Option<&str>, + limit: usize, + ) -> anyhow::Result { + let conn = self.conn()?; + let edges = crate::registry::call_graph::query_call_edges( + &conn, + repo_id, + callee.filter(|s| !s.is_empty()), + caller.filter(|s| !s.is_empty()), + file.filter(|s| !s.is_empty()), + limit, + )?; + let calls: Vec = edges + .into_iter() + .map(|e| { + serde_json::json!({ + "caller_file": e.caller_file, + "caller_symbol": e.caller_symbol, + "caller_line": e.caller_line, + "callee_name": e.callee_name, + }) + }) + .collect(); + Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "count": calls.len(), + "calls": calls + })) + } + + fn query_dependencies( + &self, + repo_id: &str, + direction: &str, + relation_type: Option<&str>, + ) -> anyhow::Result { + let conn = self.conn()?; + let rel_filter = relation_type.filter(|s| !s.is_empty()); + let label = if direction == "incoming" || direction == "reverse" { + "reverse dependencies" + } else { + "dependencies" + }; + let rows = if direction == "incoming" || direction == "reverse" { + crate::dependency_graph::list_reverse_dependencies(&conn, repo_id)? + } else { + crate::dependency_graph::list_dependencies(&conn, repo_id)? + }; + let deps: Vec = rows + .into_iter() + .filter(|(_, rel, _)| rel_filter.is_none_or(|f| f == rel)) + .map(|(id, rel, conf)| { + serde_json::json!({ + "repo_id": id, + "relation_type": rel, + "confidence": conf, + }) + }) + .collect(); + Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "direction": direction, + "label": label, + "count": deps.len(), + "dependencies": deps + })) + } + + fn query_code_symbols( + &self, + repo_id: &str, + name: Option<&str>, + symbol_type: Option<&str>, + file: Option<&str>, + limit: usize, + ) -> anyhow::Result { + let conn = self.conn()?; + let symbols = crate::registry::code_symbols::query_code_symbols( + &conn, + repo_id, + name, + symbol_type, + file, + limit, + )?; + let out: Vec = symbols + .iter() + .map(|s| { + serde_json::json!({ + "file_path": s.file_path, + "symbol_type": s.symbol_type, + "name": s.name, + "line_start": s.line_start, + "line_end": s.line_end, + "signature": s.signature, + }) + }) + .collect(); + Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "count": out.len(), + "symbols": out + })) + } + + fn query_dead_code( + &self, + repo_id: &str, + include_pub: bool, + limit: usize, + ) -> anyhow::Result { + let conn = self.conn()?; + let dead = crate::registry::dead_code::query_dead_code(&conn, repo_id, include_pub, limit)?; + let out: Vec = dead + .iter() + .map(|d| { + serde_json::json!({ + "file_path": d.file_path, + "name": d.name, + "line_start": d.line_start, + "signature": d.signature, + }) + }) + .collect(); + Ok(serde_json::json!({ + "success": true, + "repo_id": repo_id, + "count": out.len(), + "dead_functions": out + })) + } +} + #[cfg(test)] mod test_helpers; diff --git a/src/registry/code_symbols.rs b/src/registry/code_symbols.rs index d4682e8..03f5a6f 100644 --- a/src/registry/code_symbols.rs +++ b/src/registry/code_symbols.rs @@ -1,5 +1,160 @@ // SPDX-License-Identifier: MIT // Copyright (c) 2026 juice094 -// RE-EXPORT ONLY — 实现已迁移至 devbase-registry-code-symbols crate. -// 禁止在本文件中添加新代码。 -pub use devbase_registry_code_symbols::*; + +// Re-export the external crate type for backward compatibility (src/repository/symbol.rs). +pub use devbase_registry_code_symbols::CodeSymbol; + +/// A single code symbol from the `code_symbols` table (RegistryClient variant). +#[derive(Debug, Clone)] +pub struct CodeSymbolRow { + pub file_path: String, + pub symbol_type: String, + pub name: String, + pub line_start: i64, + pub line_end: i64, + pub signature: Option, +} + +/// Query code symbols for a specific repository. +pub fn query_code_symbols( + conn: &rusqlite::Connection, + repo_id: &str, + name: Option<&str>, + symbol_type: Option<&str>, + file: Option<&str>, + limit: usize, +) -> anyhow::Result> { + let mut sql = String::from( + "SELECT file_path, symbol_type, name, line_start, line_end, signature \ + FROM code_symbols WHERE repo_id = ?1", + ); + let mut params: Vec> = vec![Box::new(repo_id.to_string())]; + if let Some(ty) = symbol_type.filter(|s| !s.is_empty()) { + sql.push_str(" AND symbol_type = ?"); + sql.push_str(&(params.len() + 1).to_string()); + params.push(Box::new(ty.to_string())); + } + if let Some(n) = name.filter(|s| !s.is_empty()) { + sql.push_str(" AND name LIKE ?"); + sql.push_str(&(params.len() + 1).to_string()); + params.push(Box::new(format!("%{}%", n))); + } + if let Some(f) = file.filter(|s| !s.is_empty()) { + sql.push_str(" AND file_path LIKE ?"); + sql.push_str(&(params.len() + 1).to_string()); + params.push(Box::new(format!("%{}%", f))); + } + sql.push_str(&format!(" ORDER BY file_path, line_start LIMIT {}", limit.min(200))); + + let param_refs: Vec<&dyn rusqlite::ToSql> = params.iter().map(|p| p.as_ref()).collect(); + let mut stmt = conn.prepare(&sql)?; + let rows = stmt.query_map(rusqlite::params_from_iter(param_refs), |row| { + Ok(CodeSymbolRow { + file_path: row.get::<_, String>(0)?, + symbol_type: row.get::<_, String>(1)?, + name: row.get::<_, String>(2)?, + line_start: row.get::<_, i64>(3)?, + line_end: row.get::<_, i64>(4)?, + signature: row.get::<_, Option>(5)?, + }) + })?; + + rows.collect::, _>>().map_err(Into::into) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_query_code_symbols_no_filter() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'function', 'foo', 10, 20, 'fn foo() {}')", + [], + ) + .unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'struct', 'Bar', 30, 40, 'struct Bar;')", + [], + ) + .unwrap(); + + let rows = query_code_symbols(&conn, "r1", None, None, None, 10).unwrap(); + assert_eq!(rows.len(), 2); + assert_eq!(rows[0].name, "foo"); + assert_eq!(rows[1].name, "Bar"); + } + + #[test] + fn test_query_code_symbols_by_symbol_type() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'function', 'foo', 10, 20, 'fn foo() {}')", + [], + ) + .unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'struct', 'Bar', 30, 40, 'struct Bar;')", + [], + ) + .unwrap(); + + let rows = query_code_symbols(&conn, "r1", None, Some("function"), None, 10).unwrap(); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].name, "foo"); + } + + #[test] + fn test_query_code_symbols_by_name() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'function', 'foobar', 10, 20, 'fn foobar() {}')", + [], + ) + .unwrap(); + + let rows = query_code_symbols(&conn, "r1", Some("oob"), None, None, 10).unwrap(); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].name, "foobar"); + } + + #[test] + fn test_query_code_symbols_by_file() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/main.rs', 'function', 'foo', 10, 20, 'fn foo() {}')", + [], + ) + .unwrap(); + + let rows = query_code_symbols(&conn, "r1", None, None, Some("main"), 10).unwrap(); + assert_eq!(rows.len(), 1); + assert_eq!(rows[0].file_path, "src/main.rs"); + } + + #[test] + fn test_query_code_symbols_limit() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + for i in 0..5 { + conn.execute( + &format!( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, line_end, signature) + VALUES ('r1', 'src/lib.rs', 'function', 'f{}', {}, {}, 'fn f{}() {{}}')", + i, i, i + 1, i + ), + [], + ) + .unwrap(); + } + + let rows = query_code_symbols(&conn, "r1", None, None, None, 3).unwrap(); + assert_eq!(rows.len(), 3); + } +} diff --git a/src/registry/dead_code.rs b/src/registry/dead_code.rs index 6241e2e..d66384f 100644 --- a/src/registry/dead_code.rs +++ b/src/registry/dead_code.rs @@ -1,5 +1,183 @@ // SPDX-License-Identifier: MIT // Copyright (c) 2026 juice094 -// RE-EXPORT ONLY — 实现已迁移至 devbase-registry-dead-code crate. -// 禁止在本文件中添加新代码。 -pub use devbase_registry_dead_code::*; + +// Re-export the external crate type for backward compatibility (src/repository/symbol.rs). +pub use devbase_registry_dead_code::DeadFunction; + +/// A potentially dead function from the code symbol index. +#[derive(Debug, Clone)] +pub struct DeadCodeRow { + pub file_path: String, + pub name: String, + pub line_start: i64, + pub signature: Option, +} + +/// Query potentially dead functions for a specific repository. +pub fn query_dead_code( + conn: &rusqlite::Connection, + repo_id: &str, + include_pub: bool, + limit: usize, +) -> anyhow::Result> { + let mut sql = String::from( + "SELECT file_path, name, line_start, signature \ + FROM code_symbols cs \ + WHERE cs.repo_id = ?1 AND cs.symbol_type = 'function' \ + AND NOT EXISTS ( \ + SELECT 1 FROM code_call_graph ccg \ + WHERE ccg.repo_id = cs.repo_id AND ccg.callee_name = cs.name \ + )", + ); + if !include_pub { + sql.push_str(" AND (cs.signature IS NULL OR cs.signature NOT LIKE 'pub%fn%')"); + } + sql.push_str(" AND cs.name != 'main'"); + sql.push_str(" AND cs.name NOT LIKE 'test_%'"); + sql.push_str(" AND cs.file_path NOT LIKE '%/tests.rs' AND cs.file_path NOT LIKE '%\\tests.rs'"); + sql.push_str(" AND (cs.attributes IS NULL OR cs.attributes NOT LIKE '%#[test]%')"); + sql.push_str(&format!(" ORDER BY cs.file_path, cs.line_start LIMIT {}", limit.min(200))); + + let mut stmt = conn.prepare(&sql)?; + let rows = stmt.query_map([repo_id], |row| { + Ok(DeadCodeRow { + file_path: row.get::<_, String>(0)?, + name: row.get::<_, String>(1)?, + line_start: row.get::<_, i64>(2)?, + signature: row.get::<_, Option>(3)?, + }) + })?; + + let mut dead = Vec::new(); + for row in rows { + dead.push(row?); + } + Ok(dead) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_query_dead_code_basic() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/lib.rs', 'function', 'unused_fn', 10, 'fn unused_fn() {}')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, false, 10).unwrap(); + assert_eq!(dead.len(), 1); + assert_eq!(dead[0].name, "unused_fn"); + } + + #[test] + fn test_query_dead_code_excludes_called() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/lib.rs', 'function', 'called_fn', 10, 'fn called_fn() {}')", + [repo], + ) + .unwrap(); + conn.execute( + "INSERT INTO code_call_graph (repo_id, caller_file, caller_symbol, caller_line, callee_name) + VALUES (?1, 'src/lib.rs', 'other', 1, 'called_fn')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, false, 10).unwrap(); + assert!(dead.is_empty()); + } + + #[test] + fn test_query_dead_code_excludes_pub_when_not_include_pub() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/lib.rs', 'function', 'pub_fn', 10, 'pub fn pub_fn() {}')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, false, 10).unwrap(); + assert!(dead.is_empty()); + + let dead = query_dead_code(&conn, repo, true, 10).unwrap(); + assert_eq!(dead.len(), 1); + } + + #[test] + fn test_query_dead_code_excludes_main() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/lib.rs', 'function', 'main', 10, 'fn main() {}')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, true, 10).unwrap(); + assert!(dead.is_empty()); + } + + #[test] + fn test_query_dead_code_excludes_test_prefix() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/lib.rs', 'function', 'test_something', 10, 'fn test_something() {}')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, true, 10).unwrap(); + assert!(dead.is_empty()); + } + + #[test] + fn test_query_dead_code_excludes_tests_rs() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature) + VALUES (?1, 'src/tests.rs', 'function', 'helper', 10, 'fn helper() {}')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, true, 10).unwrap(); + assert!(dead.is_empty()); + } + + #[test] + fn test_query_dead_code_excludes_test_attribute() { + let conn = crate::registry::WorkspaceRegistry::init_in_memory().unwrap(); + let repo = "r1"; + + conn.execute( + "INSERT INTO code_symbols (repo_id, file_path, symbol_type, name, line_start, signature, attributes) + VALUES (?1, 'src/lib.rs', 'function', 'my_test', 10, 'fn my_test() {}', '#[test]')", + [repo], + ) + .unwrap(); + + let dead = query_dead_code(&conn, repo, true, 10).unwrap(); + assert!(dead.is_empty()); + } +} diff --git a/src/scan.rs b/src/scan.rs index 1bd58be..4c8522c 100644 --- a/src/scan.rs +++ b/src/scan.rs @@ -2,6 +2,7 @@ // Copyright (c) 2026 juice094 use crate::registry::repo; use crate::registry::{CodeMetrics, OplogEntry, RemoteEntry, RepoEntry}; +use crate::storage::AppContext; use chrono::Utc; use git2::Repository; use r2d2::Pool; @@ -558,6 +559,16 @@ pub fn compute_code_metrics(path: &str) -> Option { }) } +impl crate::clients::ScanClient for AppContext { + async fn scan_directory( + &self, + path: &str, + register: bool, + ) -> anyhow::Result { + crate::scan::run_json(path, register, &self.pool()).await + } +} + #[cfg(test)] mod tests { use super::*; diff --git a/src/search.rs b/src/search.rs index a6ac84b..5d3d278 100644 --- a/src/search.rs +++ b/src/search.rs @@ -3,6 +3,7 @@ // Copyright (c) 2026 juice094 pub mod hybrid; +pub mod symbol_index; use crate::storage::StorageBackend; use std::path::PathBuf; @@ -93,20 +94,11 @@ fn add_doc( tags: &[String], doc_type: &str, ) -> Result<(), TantivyError> { - // TODO(veto-audit-2026-04-26): RF-6 expect — schema 字段 lookup,不变量由 init_index 保证。可接受,但建议统一封装为辅助函数。 - let id_f = schema.get_field("id").expect("schema field 'id' defined in init_index"); - // TODO(veto-audit-2026-04-26): RF-6 expect — schema 字段 lookup,init_index 保证。 - let title_f = schema.get_field("title").expect("schema field 'title' defined in init_index"); - let content_f = schema - .get_field("content") - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - .expect("schema field 'content' defined in init_index"); - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - let tags_f = schema.get_field("tags").expect("schema field 'tags' defined in init_index"); - let doc_type_f = schema - .get_field("doc_type") - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - .expect("schema field 'doc_type' defined in init_index"); + let id_f = schema.get_field("id")?; + let title_f = schema.get_field("title")?; + let content_f = schema.get_field("content")?; + let tags_f = schema.get_field("tags")?; + let doc_type_f = schema.get_field("doc_type")?; let mut doc = TantivyDocument::default(); doc.add_text(id_f, id); @@ -124,8 +116,7 @@ pub fn delete_repo_doc( schema: &Schema, repo_id: &str, ) -> Result<(), TantivyError> { - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - let id = schema.get_field("id").expect("schema field 'id' defined in init_index"); + let id = schema.get_field("id")?; let term = tantivy::Term::from_field_text(id, repo_id); writer.delete_term(term); Ok(()) @@ -161,7 +152,7 @@ fn list_indexed_repo_ids_with_reader( ) -> Result, TantivyError> { let searcher = reader.searcher(); let schema = index.schema(); - let id_field = schema.get_field("id").expect("schema field 'id' defined in init_index"); + let id_field = schema.get_field("id")?; let all_query = tantivy::query::AllQuery; // Use a generous limit; typical deployment has < 1000 repos. @@ -264,16 +255,10 @@ fn search_with_reader( let schema = index.schema(); let searcher = reader.searcher(); - let title = schema.get_field("title").expect("schema field 'title' defined in init_index"); - let content = schema - .get_field("content") - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - .expect("schema field 'content' defined in init_index"); - let tags = schema.get_field("tags").expect("schema field 'tags' defined in init_index"); - let doc_type_f = schema - .get_field("doc_type") - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - .expect("schema field 'doc_type' defined in init_index"); + let title = schema.get_field("title")?; + let content = schema.get_field("content")?; + let tags = schema.get_field("tags")?; + let doc_type_f = schema.get_field("doc_type")?; let query_parser = QueryParser::for_index(index, vec![title, content, tags]); let text_query = query_parser.parse_query(query_str)?; @@ -294,8 +279,7 @@ fn search_with_reader( let top_docs = searcher.search(&*final_query, &TopDocs::with_limit(limit).order_by_score())?; - // TODO(veto-audit-2026-04-26): RF-6 expect — 同上。 - let id_field = schema.get_field("id").expect("schema field 'id' defined in init_index"); + let id_field = schema.get_field("id")?; let mut results = Vec::new(); for (score, doc_address) in top_docs { let doc: TantivyDocument = searcher.doc(doc_address)?; diff --git a/src/search/hybrid.rs b/src/search/hybrid.rs index 47fa75a..20ab7f9 100644 --- a/src/search/hybrid.rs +++ b/src/search/hybrid.rs @@ -13,28 +13,40 @@ use std::collections::HashMap; use crate::semantic_index::SemanticSearchRow; -/// Keyword search over code symbols using SQLite LIKE on name and signature. +/// Keyword search over code symbols. /// -/// Score heuristic: -/// - name match = 3.0 -/// - signature match = 1.0 -/// -/// The query string is whitespace-split and each token contributes -/// independently. +/// Primary path: Tantivy BM25 via symbol_index. +/// Fallback path: SQLite LIKE (for repos without a symbol index). pub fn keyword_search_symbols( conn: &rusqlite::Connection, repo_id: &str, query: &str, limit: usize, +) -> anyhow::Result> { + // Try Tantivy BM25 first + match crate::search::symbol_index::search_symbols(query, limit, Some(repo_id)) { + Ok(results) if !results.is_empty() => return Ok(results), + Ok(_) => {} // empty results, try fallback + Err(e) => { + tracing::debug!("Symbol index search failed for {}: {}", repo_id, e); + } + } + + // Fallback: SQLite LIKE (for repos without symbol index or when index is empty) + keyword_search_symbols_fallback(conn, repo_id, query, limit) +} + +fn keyword_search_symbols_fallback( + conn: &rusqlite::Connection, + repo_id: &str, + query: &str, + limit: usize, ) -> anyhow::Result> { let tokens: Vec<&str> = query.split_whitespace().collect(); if tokens.is_empty() { return Ok(Vec::new()); } - // Simple per-token query with aggregated scoring. - // For each token we query once and accumulate scores in a HashMap. - // name match = 3.0, signature match = 1.0. let mut accum: HashMap = HashMap::new(); for token in &tokens { @@ -85,13 +97,12 @@ pub fn keyword_search_symbols( /// fused ranking. The standard k constant is 60.0. /// /// Items are deduplicated by `(repo_id, name, file_path)`. -pub fn rrf_merge(lists: Vec>, k: f32) -> Vec { +pub fn rrf_merge(mut lists: Vec>, k: f32) -> Vec { if lists.is_empty() { return Vec::new(); } if lists.len() == 1 { - // TODO(veto-audit-2026-04-26): RF-6 expect — 前置 len==1 检查已做,风险低。可保留或改为 `unwrap_or_default`。 - return lists.into_iter().next().expect("lists len == 1 checked above"); + return lists.remove(0); } let mut accum: HashMap = HashMap::new(); @@ -153,14 +164,7 @@ pub fn hybrid_search_symbols( match lists.len() { 0 => Ok(Vec::new()), - 1 => Ok(lists - .into_iter() - .next() - // TODO(veto-audit-2026-04-26): RF-6 expect — 前置 len==1 检查已做,风险低。 - .expect("lists len == 1 checked above") - .into_iter() - .take(limit) - .collect()), + 1 => Ok(lists.remove(0).into_iter().take(limit).collect()), _ => { let merged = rrf_merge(lists, 60.0); Ok(merged.into_iter().take(limit).collect()) diff --git a/src/search/symbol_index.rs b/src/search/symbol_index.rs new file mode 100644 index 0000000..d198bbb --- /dev/null +++ b/src/search/symbol_index.rs @@ -0,0 +1,363 @@ +// SPDX-License-Identifier: MIT +// Copyright (c) 2026 juice094 +//! Tantivy-based BM25 search index for code symbols. +//! +//! Replaces the SQLite `LIKE` fallback in `hybrid.rs` with proper +//! full-text retrieval over symbol names and signatures. + +use crate::storage::StorageBackend; +use std::path::Path; +use tantivy::{ + Index, IndexReader, IndexWriter, ReloadPolicy, TantivyDocument, TantivyError, + collector::TopDocs, + query::QueryParser, + schema::{STORED, Schema, TEXT, Value}, +}; + +const SYMBOL_INDEX_DIR: &str = "symbol_index"; + +fn symbol_index_path() -> Result { + crate::storage::DefaultStorageBackend {} + .symbol_index_path() + .map_err(|e| TantivyError::InvalidArgument(e.to_string())) +} + +fn build_schema() -> Schema { + let mut sb = Schema::builder(); + sb.add_text_field("repo_id", TEXT | STORED); + sb.add_text_field("name", TEXT | STORED); + sb.add_text_field("signature", TEXT | STORED); + sb.add_text_field("file_path", TEXT | STORED); + sb.add_text_field("line_start", STORED); + sb.build() +} + +/// Initialize the symbol index (uses default storage backend). +pub fn init_index() -> Result<(Index, IndexReader), TantivyError> { + let path = symbol_index_path()?; + init_index_at(&path) +} + +/// Initialize at an explicit path (tests + hermetic isolation). +pub fn init_index_at(path: &Path) -> Result<(Index, IndexReader), TantivyError> { + std::fs::create_dir_all(path)?; + let schema = build_schema(); + let index = match Index::open_in_dir(path) { + Ok(idx) => { + if idx.schema() == schema { + idx + } else { + drop(idx); + let _ = std::fs::remove_dir_all(path); + std::fs::create_dir_all(path)?; + Index::create_in_dir(path, schema)? + } + } + Err(_) => Index::create_in_dir(path, schema)?, + }; + let reader = index.reader_builder().reload_policy(ReloadPolicy::Manual).try_into()?; + Ok((index, reader)) +} + +pub fn get_writer(index: &Index) -> Result { + index.writer(50_000_000) +} + +/// Add a single code symbol document. +pub fn add_symbol_doc( + writer: &mut IndexWriter, + schema: &Schema, + repo_id: &str, + name: &str, + signature: Option<&str>, + file_path: &str, + line_start: usize, +) -> Result<(), TantivyError> { + let repo_f = schema.get_field("repo_id")?; + let name_f = schema.get_field("name")?; + let sig_f = schema.get_field("signature")?; + let path_f = schema.get_field("file_path")?; + let line_f = schema.get_field("line_start")?; + + let mut doc = TantivyDocument::default(); + doc.add_text(repo_f, repo_id); + doc.add_text(name_f, name); + if let Some(s) = signature { + doc.add_text(sig_f, s); + } + doc.add_text(path_f, file_path); + doc.add_text(line_f, &line_start.to_string()); + writer.add_document(doc)?; + Ok(()) +} + +/// Bulk-add symbols for a repo (deletes existing repo symbols first). +pub fn add_symbols( + writer: &mut IndexWriter, + schema: &Schema, + repo_id: &str, + symbols: &[crate::semantic_index::CodeSymbol], +) -> Result<(), TantivyError> { + // Delete existing symbols for this repo + delete_repo_symbols(writer, schema, repo_id)?; + for sym in symbols { + add_symbol_doc( + writer, + schema, + repo_id, + &sym.name, + sym.signature.as_deref(), + &sym.file_path.to_string_lossy(), + sym.line_start, + )?; + } + Ok(()) +} + +/// Delete all symbols belonging to a repo. +pub fn delete_repo_symbols( + writer: &mut IndexWriter, + schema: &Schema, + repo_id: &str, +) -> Result<(), TantivyError> { + let repo_f = schema.get_field("repo_id")?; + let term = tantivy::Term::from_field_text(repo_f, repo_id); + writer.delete_term(term); + Ok(()) +} + +pub fn commit_writer(writer: &mut IndexWriter) -> Result<(), TantivyError> { + writer.commit()?; + Ok(()) +} + +/// BM25 search over code symbols, optionally filtered by repo_id. +/// +/// Queries the `name` and `signature` fields. Returns +/// Vec<(repo_id, name, file_path, line_start, bm25_score)>. +pub fn search_symbols( + query_str: &str, + limit: usize, + repo_id: Option<&str>, +) -> Result, TantivyError> { + let path = symbol_index_path()?; + search_symbols_at(&path, query_str, limit, repo_id) +} + +/// Search at an explicit path. +pub fn search_symbols_at( + path: &Path, + query_str: &str, + limit: usize, + repo_id: Option<&str>, +) -> Result, TantivyError> { + let (index, reader) = init_index_at(path)?; + let schema = index.schema(); + let searcher = reader.searcher(); + + let name_f = schema.get_field("name")?; + let sig_f = schema.get_field("signature")?; + let repo_f = schema.get_field("repo_id")?; + let path_f = schema.get_field("file_path")?; + let line_f = schema.get_field("line_start")?; + + let parser = QueryParser::for_index(&index, vec![name_f, sig_f]); + let text_query = parser.parse_query(query_str)?; + + // Build combined query: text_query AND repo_id:filter (if specified) + let final_query: Box = if let Some(rid) = repo_id { + let repo_term_query = tantivy::query::TermQuery::new( + tantivy::Term::from_field_text(repo_f, rid), + tantivy::schema::IndexRecordOption::Basic, + ); + Box::new(tantivy::query::BooleanQuery::new(vec![ + (tantivy::query::Occur::Must, text_query), + (tantivy::query::Occur::Must, Box::new(repo_term_query)), + ])) + } else { + text_query + }; + + let top_docs = searcher.search(&*final_query, &TopDocs::with_limit(limit).order_by_score())?; + + let mut results = Vec::new(); + for (score, doc_addr) in top_docs { + let doc: TantivyDocument = searcher.doc(doc_addr)?; + let repo_id = doc.get_first(repo_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let name = doc.get_first(name_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let file_path = doc.get_first(path_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let line_start: i64 = doc + .get_first(line_f) + .and_then(|v| v.as_str()) + .unwrap_or("0") + .parse() + .unwrap_or(0); + results.push((repo_id, name, file_path, line_start, score)); + } + Ok(results) +} + +/// List all repo IDs in the symbol index. +pub fn list_indexed_repo_ids() -> Result, TantivyError> { + let path = symbol_index_path()?; + list_indexed_repo_ids_at(&path) +} + +pub fn list_indexed_repo_ids_at(path: &Path) -> Result, TantivyError> { + let (index, reader) = init_index_at(path)?; + let searcher = reader.searcher(); + let schema = index.schema(); + let repo_f = schema.get_field("repo_id")?; + + let all_query = tantivy::query::AllQuery; + let top_docs = searcher.search(&all_query, &TopDocs::with_limit(10_000).order_by_score())?; + + let mut ids = Vec::new(); + for (_score, doc_addr) in top_docs { + let doc: TantivyDocument = searcher.doc(doc_addr)?; + if let Some(id) = doc.get_first(repo_f).and_then(|v| v.as_str()) { + ids.push(id.to_string()); + } + } + ids.sort_unstable(); + ids.dedup(); + Ok(ids) +} + +#[cfg(test)] +pub(crate) static SYMBOL_INDEX_TEST_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(()); + +#[cfg(test)] +mod tests { + use super::*; + + fn with_temp_symbol_index(f: F) + where + F: FnOnce(&Index, &Schema, &mut IndexWriter), + { + let _guard = SYMBOL_INDEX_TEST_LOCK.lock().unwrap_or_else(|p| p.into_inner()); + let schema = build_schema(); + let idx = Index::create_in_ram(schema.clone()); + let mut writer = idx.writer(15_000_000).unwrap(); + f(&idx, &schema, &mut writer); + drop(writer); + drop(idx); + } + + #[test] + fn test_add_and_search_symbol() { + with_temp_symbol_index(|idx, schema, writer| { + add_symbol_doc( + writer, + schema, + "repo1", + "handle_error", + Some("pub fn handle_error(e: Error)"), + "src/lib.rs", + 42, + ) + .unwrap(); + writer.commit().unwrap(); + + let reader = idx.reader().unwrap(); + let results = search_with_reader(&reader, idx, "handle", 10, Some("repo1")).unwrap(); + assert_eq!(results.len(), 1); + assert_eq!(results[0].1, "handle_error"); + assert_eq!(results[0].3, 42); + }); + } + + #[test] + fn test_delete_repo_symbols() { + with_temp_symbol_index(|idx, schema, writer| { + add_symbol_doc(writer, schema, "repo1", "foo", None, "a.rs", 1).unwrap(); + add_symbol_doc(writer, schema, "repo2", "bar", None, "b.rs", 2).unwrap(); + writer.commit().unwrap(); + + delete_repo_symbols(writer, schema, "repo1").unwrap(); + writer.commit().unwrap(); + + let reader = idx.reader().unwrap(); + let results = search_with_reader(&reader, idx, "foo", 10, Some("repo1")).unwrap(); + assert!(results.is_empty()); + + let results = search_with_reader(&reader, idx, "bar", 10, Some("repo2")).unwrap(); + assert_eq!(results.len(), 1); + }); + } + + #[test] + fn test_search_signature_match() { + with_temp_symbol_index(|idx, schema, writer| { + add_symbol_doc( + writer, + schema, + "repo1", + "authenticate", + Some("pub fn authenticate(token: &str)"), + "src/auth.rs", + 10, + ) + .unwrap(); + writer.commit().unwrap(); + + let reader = idx.reader().unwrap(); + let results = search_with_reader(&reader, idx, "token", 10, Some("repo1")).unwrap(); + assert_eq!(results.len(), 1); + assert_eq!(results[0].1, "authenticate"); + }); + } + + // Helper that works with an existing reader for in-RAM tests. + fn search_with_reader( + reader: &IndexReader, + index: &Index, + query: &str, + limit: usize, + repo_id: Option<&str>, + ) -> Result, TantivyError> { + let schema = index.schema(); + let searcher = reader.searcher(); + let name_f = schema.get_field("name").unwrap(); + let sig_f = schema.get_field("signature").unwrap(); + let repo_f = schema.get_field("repo_id").unwrap(); + let path_f = schema.get_field("file_path").unwrap(); + let line_f = schema.get_field("line_start").unwrap(); + + let parser = QueryParser::for_index(index, vec![name_f, sig_f]); + let text_query = parser.parse_query(query).unwrap(); + + let final_query: Box = if let Some(rid) = repo_id { + let repo_term = tantivy::query::TermQuery::new( + tantivy::Term::from_field_text(repo_f, rid), + tantivy::schema::IndexRecordOption::Basic, + ); + Box::new(tantivy::query::BooleanQuery::new(vec![ + (tantivy::query::Occur::Must, text_query), + (tantivy::query::Occur::Must, Box::new(repo_term)), + ])) + } else { + text_query + }; + + let top_docs = + searcher.search(&*final_query, &TopDocs::with_limit(limit).order_by_score())?; + + let mut results = Vec::new(); + for (score, doc_addr) in top_docs { + let doc: TantivyDocument = searcher.doc(doc_addr)?; + let repo_id = doc.get_first(repo_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let name = doc.get_first(name_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let file_path = + doc.get_first(path_f).and_then(|v| v.as_str()).unwrap_or("").to_string(); + let line_start: i64 = doc + .get_first(line_f) + .and_then(|v| v.as_str()) + .unwrap_or("0") + .parse() + .unwrap_or(0); + results.push((repo_id, name, file_path, line_start, score)); + } + Ok(results) + } +} diff --git a/src/semantic_index/call_graph.rs b/src/semantic_index/call_graph.rs index bd73b87..7ae8d80 100644 --- a/src/semantic_index/call_graph.rs +++ b/src/semantic_index/call_graph.rs @@ -35,8 +35,18 @@ pub fn extract_calls_from_file(file_path: &Path, source: &str) -> Vec } }; + extract_calls_from_tree(&tree, file_path, source.as_bytes(), lang) +} + +/// Extract calls from an already-parsed tree-sitter tree. +/// This avoids re-parsing when both symbols and calls are needed. +pub(crate) fn extract_calls_from_tree( + tree: &tree_sitter::Tree, + file_path: &Path, + source_bytes: &[u8], + lang: Lang, +) -> Vec { let mut calls = Vec::new(); - let source_bytes = source.as_bytes(); let root = tree.root_node(); let mut cursor = root.walk(); walk_tree_for_calls(&mut cursor, file_path, source_bytes, lang, &mut calls, None); @@ -84,6 +94,7 @@ fn extract_current_function_name( lang: Lang, ) -> Option { match lang { + #[cfg(feature = "lang-rust")] Lang::Rust => { if node.kind() == "function_item" || node.kind() == "closure_expression" { extract_node_name(node, source_bytes) @@ -91,6 +102,7 @@ fn extract_current_function_name( None } } + #[cfg(feature = "lang-python")] Lang::Python => { if node.kind() == "function_definition" { extract_node_name(node, source_bytes) @@ -98,6 +110,7 @@ fn extract_current_function_name( None } } + #[cfg(feature = "lang-js-ts")] Lang::JsTs => { if node.kind() == "function_declaration" || node.kind() == "method_definition" { if node.kind() == "method_definition" { @@ -109,6 +122,7 @@ fn extract_current_function_name( None } } + #[cfg(feature = "lang-go")] Lang::Go => { if node.kind() == "function_declaration" { extract_node_name(node, source_bytes) @@ -127,6 +141,7 @@ fn extract_callee_name( lang: Lang, ) -> Option { match lang { + #[cfg(feature = "lang-rust")] Lang::Rust => { if node.kind() == "call_expression" { let func_node = node.child(0)?; @@ -137,6 +152,7 @@ fn extract_callee_name( None } } + #[cfg(feature = "lang-python")] Lang::Python => { if node.kind() == "call" { let func_node = node.child(0)?; @@ -145,6 +161,7 @@ fn extract_callee_name( None } } + #[cfg(feature = "lang-js-ts")] Lang::JsTs => { if node.kind() == "call_expression" { let func_node = node.child(0)?; @@ -153,6 +170,7 @@ fn extract_callee_name( None } } + #[cfg(feature = "lang-go")] Lang::Go => { if node.kind() == "call_expression" { let func_node = node.child(0)?; diff --git a/src/semantic_index/mod.rs b/src/semantic_index/mod.rs index 5e44dde..ba1d61a 100644 --- a/src/semantic_index/mod.rs +++ b/src/semantic_index/mod.rs @@ -85,20 +85,29 @@ pub type SemanticSearchRow = (String, String, String, i64, f32); // --------------------------------------------------------------------------- #[derive(Clone, Copy)] -enum Lang { +pub(crate) enum Lang { + #[cfg(feature = "lang-rust")] Rust, + #[cfg(feature = "lang-python")] Python, + #[cfg(feature = "lang-js-ts")] JsTs, + #[cfg(feature = "lang-go")] Go, } impl Lang { fn from_ext(ext: &str) -> Option { match ext { + #[cfg(feature = "lang-rust")] "rs" => Some(Lang::Rust), + #[cfg(feature = "lang-python")] "py" => Some(Lang::Python), + #[cfg(feature = "lang-js-ts")] "js" | "ts" | "jsx" => Some(Lang::JsTs), + #[cfg(feature = "lang-js-ts")] "tsx" => Some(Lang::JsTs), + #[cfg(feature = "lang-go")] "go" => Some(Lang::Go), _ => None, } @@ -106,9 +115,13 @@ impl Lang { fn parser_language(self) -> tree_sitter::Language { match self { + #[cfg(feature = "lang-rust")] Lang::Rust => tree_sitter_rust::LANGUAGE.into(), + #[cfg(feature = "lang-python")] Lang::Python => tree_sitter_python::LANGUAGE.into(), + #[cfg(feature = "lang-js-ts")] Lang::JsTs => tree_sitter_typescript::LANGUAGE_TYPESCRIPT.into(), + #[cfg(feature = "lang-go")] Lang::Go => tree_sitter_go::LANGUAGE.into(), } } @@ -132,7 +145,7 @@ pub fn should_skip_dir(path: &Path, exclude: &[String]) -> bool { /// Extract symbols from a single source file. /// /// Supports Rust, Python, JavaScript/TypeScript, and Go. -pub fn index_repo_full(repo_path: &Path) -> (Vec, Vec) { +pub fn index_repo_full(repo_path: &Path) -> anyhow::Result<(Vec, Vec)> { let exts: &[&str] = &["rs", "py", "js", "ts", "jsx", "tsx", "go"]; let cfg = crate::config::Config::load().unwrap_or_default(); let mut exclude = cfg.scan.exclude_patterns; @@ -169,7 +182,7 @@ pub fn index_repo_full(repo_path: &Path) -> (Vec, Vec) { for path in files { process_file(repo_path, &path, &mut all_symbols, &mut all_calls); } - return (all_symbols, all_calls); + return Ok((all_symbols, all_calls)); } std::thread::scope(|s| { @@ -177,29 +190,28 @@ pub fn index_repo_full(repo_path: &Path) -> (Vec, Vec) { let mut handles = Vec::with_capacity(num_threads); for chunk in files.chunks(chunk_size) { - handles.push( - std::thread::Builder::new() - .stack_size(4 * 1024 * 1024) - .spawn_scoped(s, move || { - let mut symbols = Vec::new(); - let mut calls = Vec::new(); - for path in chunk { - process_file(repo_path, path, &mut symbols, &mut calls); - } - (symbols, calls) - }) - .expect("failed to spawn index worker"), - ); + handles.push(std::thread::Builder::new().stack_size(4 * 1024 * 1024).spawn_scoped( + s, + move || { + let mut symbols = Vec::new(); + let mut calls = Vec::new(); + for path in chunk { + process_file(repo_path, path, &mut symbols, &mut calls); + } + (symbols, calls) + }, + )?); } let mut all_symbols = Vec::new(); let mut all_calls = Vec::new(); for handle in handles { - let (s, c) = handle.join().unwrap(); + let (s, c) = + handle.join().map_err(|e| anyhow::anyhow!("index worker panicked: {:?}", e))?; all_symbols.extend(s); all_calls.extend(c); } - (all_symbols, all_calls) + Ok((all_symbols, all_calls)) }) } @@ -233,10 +245,34 @@ fn process_file( match std::fs::read_to_string(path) { Ok(source) => { let rel_path = path.strip_prefix(repo_path).unwrap_or(path); - let symbols = extract_symbols(rel_path, &source); - all_symbols.extend(symbols); - let calls = extract_calls_from_file(rel_path, &source); + let ext = rel_path.extension().and_then(|e| e.to_str()); + let lang = match ext { + Some(ext) => match Lang::from_ext(ext) { + Some(l) => l, + None => return, + }, + None => return, + }; + + let mut parser = tree_sitter::Parser::new(); + let ts_lang = lang.parser_language(); + if let Err(e) = parser.set_language(&ts_lang) { + warn!("Failed to set tree-sitter language: {}", e); + return; + } + let tree = match parser.parse(&source, None) { + Some(t) => t, + None => { + warn!("Failed to parse {:?}", rel_path); + return; + } + }; + + let source_bytes = source.as_bytes(); + let symbols = extract_symbols_from_tree(&tree, rel_path, source_bytes, lang); + all_symbols.extend(symbols); + let calls = extract_calls_from_tree(&tree, rel_path, source_bytes, lang); all_calls.extend(calls); } Err(e) => { @@ -246,8 +282,8 @@ fn process_file( } /// Scan a repository for source files and extract all symbols. -pub fn index_repo(repo_path: &Path) -> Vec { - index_repo_full(repo_path).0 +pub fn index_repo(repo_path: &Path) -> anyhow::Result> { + Ok(index_repo_full(repo_path)?.0) } #[cfg(test)] @@ -575,7 +611,7 @@ class MyClass: ) .unwrap(); - let (symbols, calls) = index_repo_full(tmp.path()); + let (symbols, calls) = index_repo_full(tmp.path()).unwrap(); assert!(!symbols.is_empty()); let names: Vec<_> = symbols.iter().map(|s| s.name.as_str()).collect(); diff --git a/src/semantic_index/symbol.rs b/src/semantic_index/symbol.rs index e69100b..416e982 100644 --- a/src/semantic_index/symbol.rs +++ b/src/semantic_index/symbol.rs @@ -37,8 +37,18 @@ fn extract_symbols_with_parser(file_path: &Path, source: &str, lang: Lang) -> Ve } }; + extract_symbols_from_tree(&tree, file_path, source.as_bytes(), lang) +} + +/// Extract symbols from an already-parsed tree-sitter tree. +/// This avoids re-parsing when both symbols and calls are needed. +pub(crate) fn extract_symbols_from_tree( + tree: &tree_sitter::Tree, + file_path: &Path, + source_bytes: &[u8], + lang: Lang, +) -> Vec { let mut symbols = Vec::new(); - let source_bytes = source.as_bytes(); let root = tree.root_node(); collect_symbols_from_node(&root, file_path, source_bytes, lang, &mut symbols); symbols @@ -72,9 +82,13 @@ fn node_to_symbol( lang: Lang, ) -> Option { match lang { + #[cfg(feature = "lang-rust")] Lang::Rust => rust_node_to_symbol(node, file_path, source_bytes), + #[cfg(feature = "lang-python")] Lang::Python => python_node_to_symbol(node, file_path, source_bytes), + #[cfg(feature = "lang-js-ts")] Lang::JsTs => js_node_to_symbol(node, file_path, source_bytes), + #[cfg(feature = "lang-go")] Lang::Go => go_node_to_symbol(node, file_path, source_bytes), } } @@ -362,6 +376,7 @@ mod tests { use super::*; #[test] + #[cfg(feature = "lang-rust")] fn test_extract_rust_attributes() { let source = r#" #[tokio::test] diff --git a/src/storage.rs b/src/storage.rs index 8822965..39fa718 100644 --- a/src/storage.rs +++ b/src/storage.rs @@ -4,6 +4,35 @@ use r2d2::Pool; use r2d2_sqlite::SqliteConnectionManager; use std::path::PathBuf; use std::sync::Arc; +use std::time::{Duration, Instant}; + +/// Session-level cache for environment tool versions. +/// Avoids spawning subprocesses on every health check. +#[derive(Debug, Clone, Default)] +pub struct EnvVersionCache { + pub rustc: Option, + pub cargo: Option, + pub node: Option, + pub go: Option, + pub cmake: Option, + pub python: Option, + pub bun: Option, + pub zig: Option, + pub java: Option, + pub fetched_at: Option, +} + +impl EnvVersionCache { + const TTL: Duration = Duration::from_secs(30); + + pub fn is_fresh(&self) -> bool { + self.fetched_at.map(|t| t.elapsed() < Self::TTL).unwrap_or(false) + } + + pub fn clear(&mut self) { + *self = Self::default(); + } +} /// 抽象数据存储后端,解耦具体路径实现。 /// @@ -19,6 +48,9 @@ pub trait StorageBackend: Send + Sync { /// Tantivy 搜索索引目录。 fn index_path(&self) -> anyhow::Result; + /// Tantivy 代码符号搜索索引目录。 + fn symbol_index_path(&self) -> anyhow::Result; + /// 自动备份目录。 fn backup_dir(&self) -> anyhow::Result; } @@ -62,6 +94,12 @@ impl StorageBackend for DefaultStorageBackend { Ok(dir.join("search_index")) } + fn symbol_index_path(&self) -> anyhow::Result { + let dir = self.data_base()?; + std::fs::create_dir_all(&dir)?; + Ok(dir.join("symbol_index")) + } + fn backup_dir(&self) -> anyhow::Result { let dir = self.data_base()?; let backup = dir.join("backups"); @@ -79,6 +117,7 @@ pub struct AppContext { pub config: crate::config::Config, pub i18n: crate::i18n::I18n, pool: Pool, + env_cache: std::sync::Mutex, } impl AppContext { @@ -98,7 +137,13 @@ impl AppContext { let pool = Self::build_pool(&path)?; let config = crate::config::Config::load()?; let i18n = crate::i18n::from_language(&config.general.language); - Ok(Self { storage, config, i18n, pool }) + Ok(Self { + storage, + config, + i18n, + pool, + env_cache: std::sync::Mutex::new(EnvVersionCache::default()), + }) } /// 使用自定义存储后端创建上下文(主要用于测试)。 @@ -115,7 +160,13 @@ impl AppContext { let pool = Self::build_pool(&path)?; let config = crate::config::Config::load()?; let i18n = crate::i18n::from_language(&config.general.language); - Ok(Self { storage, config, i18n, pool }) + Ok(Self { + storage, + config, + i18n, + pool, + env_cache: std::sync::Mutex::new(EnvVersionCache::default()), + }) } fn build_pool(path: &std::path::Path) -> anyhow::Result> { @@ -140,17 +191,51 @@ impl AppContext { pub fn pool(&self) -> Pool { self.pool.clone() } + + /// 获取环境版本缓存的只读快照。 + pub fn env_cache(&self) -> anyhow::Result { + let guard = self + .env_cache + .lock() + .map_err(|e| anyhow::anyhow!("env_cache poisoned: {}", e))?; + Ok(guard.clone()) + } + + /// 更新环境版本缓存。 + pub fn set_env_cache(&self, cache: EnvVersionCache) -> anyhow::Result<()> { + let mut guard = self + .env_cache + .lock() + .map_err(|e| anyhow::anyhow!("env_cache poisoned: {}", e))?; + *guard = cache; + Ok(()) + } +} + +/// Result of a startup consistency scan. +#[allow(dead_code)] +pub(crate) struct RepairResult { + /// Tantivy documents whose repo no longer exists in SQLite. + pub orphans: usize, + /// SQLite entities that are missing from the Tantivy index. + pub missing_from_index: usize, } /// Startup consistency scan: detect Tantivy documents whose repo no longer exists in SQLite. +/// Also detects SQLite repos that are missing from the Tantivy index. /// Inserts orphan records into `orphan_tantivy_docs` for lazy cleanup during next index. -pub(crate) fn repair_tantivy_consistency(conn: &mut rusqlite::Connection) -> anyhow::Result { +pub(crate) fn repair_tantivy_consistency( + conn: &mut rusqlite::Connection, +) -> anyhow::Result { let backend = crate::storage::DefaultStorageBackend {}; let index_path = match backend.index_path() { Ok(p) => p, Err(e) => { tracing::warn!("Failed to resolve index path: {}", e); - return Ok(0); + return Ok(RepairResult { + orphans: 0, + missing_from_index: 0, + }); } }; repair_tantivy_consistency_at(&index_path, conn) @@ -160,14 +245,18 @@ pub(crate) fn repair_tantivy_consistency(conn: &mut rusqlite::Connection) -> any pub(crate) fn repair_tantivy_consistency_at( index_path: &std::path::Path, conn: &mut rusqlite::Connection, -) -> anyhow::Result { - let tantivy_ids = match crate::search::list_indexed_repo_ids_at(index_path) { - Ok(ids) => ids, - Err(e) => { - tracing::warn!("Failed to list Tantivy repo IDs: {}", e); - return Ok(0); - } - }; +) -> anyhow::Result { + let tantivy_ids: std::collections::HashSet = + match crate::search::list_indexed_repo_ids_at(index_path) { + Ok(ids) => ids.into_iter().collect(), + Err(e) => { + tracing::warn!("Failed to list Tantivy repo IDs: {}", e); + return Ok(RepairResult { + orphans: 0, + missing_from_index: 0, + }); + } + }; let sqlite_ids: std::collections::HashSet = { let mut stmt = conn.prepare("SELECT id FROM entities WHERE entity_type = ?1")?; @@ -177,7 +266,6 @@ pub(crate) fn repair_tantivy_consistency_at( }; // Clear stale orphans: repos that are now present in SQLite but still marked orphan - let mut orphaned = 0usize; { let mut stmt = conn.prepare("SELECT repo_id FROM orphan_tantivy_docs")?; let rows = stmt.query_map([], |row| row.get::<_, String>(0))?; @@ -189,6 +277,7 @@ pub(crate) fn repair_tantivy_consistency_at( } // Record new orphans: Tantivy has doc but SQLite has no entity + let mut orphaned = 0usize; for repo_id in &tantivy_ids { if !sqlite_ids.contains(repo_id) { conn.execute( @@ -202,425 +291,23 @@ pub(crate) fn repair_tantivy_consistency_at( if orphaned > 0 { tracing::info!("Detected {} orphan Tantivy document(s)", orphaned); } - Ok(orphaned) -} - -impl crate::clients::ScanClient for AppContext { - async fn scan_directory( - &self, - path: &str, - register: bool, - ) -> anyhow::Result { - crate::scan::run_json(path, register, &self.pool()).await - } -} - -impl crate::clients::HealthClient for AppContext { - async fn check_health(&self, detail: bool) -> anyhow::Result { - let conn = self.conn()?; - crate::health::run_json(&conn, detail, 0, 1, self.config.cache.ttl_seconds, &self.i18n) - .await - } -} - -impl crate::clients::SyncClient for AppContext { - async fn sync_repos( - &self, - dry_run: bool, - filter_tags: Option>, - ) -> anyhow::Result { - let conn = self.conn()?; - let filter_tags_str = filter_tags.as_deref().map(|v| v.join(",")); - crate::sync::run_json(&conn, dry_run, filter_tags_str.as_deref(), None, &self.i18n).await - } -} - -impl crate::clients::DigestClient for AppContext { - fn generate_daily_digest(&self) -> anyhow::Result { - let conn = self.conn()?; - let text = crate::digest::generate_daily_digest(&conn, &self.config, &self.i18n)?; - Ok(serde_json::json!({ "success": true, "digest": text })) - } -} - -impl crate::clients::KnowledgeClient for AppContext { - fn run_index(&self, path: &str) -> anyhow::Result { - let mut conn = self.conn()?; - let count = crate::knowledge_engine::run_index(&mut conn, path)?; - Ok(serde_json::json!({ "success": true, "indexed": count, "errors": 0 })) - } - - fn save_note( - &self, - repo_id: &str, - text: &str, - author: &str, - ) -> anyhow::Result { - let conn = self.conn()?; - crate::registry::knowledge::save_note(&conn, repo_id, text, author)?; - Ok(serde_json::json!({ "success": true })) - } - - fn save_summary( - &self, - repo_id: &str, - desc: &str, - author: &str, - ) -> anyhow::Result { - let conn = self.conn()?; - crate::registry::knowledge::save_summary(&conn, repo_id, desc, author)?; - Ok(serde_json::json!({ "success": true })) - } - fn get_paper(&self, arxiv_id: &str) -> anyhow::Result { - let conn = self.conn()?; - let papers = crate::registry::knowledge::list_papers(&conn)?; - match papers.into_iter().find(|p| p.id == arxiv_id) { - Some(p) => Ok(serde_json::json!({ - "success": true, - "id": p.id, - "title": p.title, - "venue": p.venue, - "year": p.year, - "pdf_path": p.pdf_path, - "tags": p.tags, - })), - None => Ok(serde_json::json!({ "success": false, "error": "Paper not found" })), + // Reverse check: SQLite entities missing from Tantivy + let mut missing = 0usize; + for repo_id in &sqlite_ids { + if !tantivy_ids.contains(repo_id) { + tracing::warn!( + "repo {} exists in SQLite but missing from Tantivy index; needs re-index", + repo_id + ); + missing += 1; } } -} -impl crate::clients::RegistryClient for AppContext { - fn list_repos(&self, _filter: Option<&str>) -> anyhow::Result { - let conn = self.conn()?; - let repos = crate::registry::repo::list_repos(&conn)?; - let results: Vec = repos - .into_iter() - .map(|r| { - serde_json::json!({ - "id": r.id, - "local_path": r.local_path, - "language": r.language, - "tags": r.tags, - "workspace_type": r.workspace_type, - "data_tier": r.data_tier, - }) - }) - .collect(); - Ok(serde_json::json!({ "success": true, "count": results.len(), "repos": results })) - } - - fn get_repo(&self, repo_id: &str) -> anyhow::Result { - let conn = self.conn()?; - let repos = crate::registry::repo::list_repos(&conn)?; - match repos.into_iter().find(|r| r.id == repo_id) { - Some(r) => Ok(serde_json::json!({ - "success": true, - "id": r.id, - "local_path": r.local_path, - "language": r.language, - "tags": r.tags, - "workspace_type": r.workspace_type, - "data_tier": r.data_tier, - })), - None => Ok(serde_json::json!({ "success": false, "error": "repo not found" })), - } - } - - fn list_modules(&self, repo_id: &str) -> anyhow::Result { - let conn = self.conn()?; - let modules = crate::registry::knowledge::list_modules(&conn, repo_id)?; - let results: Vec = modules - .into_iter() - .map(|(name, ty, path)| { - serde_json::json!({ - "name": name, - "type": ty, - "path": path, - }) - }) - .collect(); - Ok(serde_json::json!({ "success": true, "count": results.len(), "modules": results })) - } - - fn save_paper(&self, paper: &serde_json::Value) -> anyhow::Result { - let conn = self.conn()?; - let paper_entry: crate::registry::PaperEntry = serde_json::from_value(paper.clone())?; - crate::registry::knowledge::save_paper(&conn, &paper_entry)?; - Ok(serde_json::json!({ "success": true })) - } - - fn save_experiment(&self, exp: &serde_json::Value) -> anyhow::Result { - let conn = self.conn()?; - let exp_entry: crate::registry::ExperimentEntry = serde_json::from_value(exp.clone())?; - crate::registry::WorkspaceRegistry::save_experiment(&conn, &exp_entry)?; - Ok(serde_json::json!({ "success": true })) - } - - fn list_code_metrics(&self) -> anyhow::Result { - let conn = self.conn()?; - let metrics = crate::registry::metrics::list_code_metrics(&conn)?; - let repos: Vec = metrics - .into_iter() - .map(|(id, m)| { - serde_json::json!({ - "repo_id": id, - "total_lines": m.total_lines, - "source_lines": m.source_lines, - "test_lines": m.test_lines, - "comment_lines": m.comment_lines, - "file_count": m.file_count, - "language_breakdown": m.language_breakdown, - "updated_at": m.updated_at.to_rfc3339() - }) - }) - .collect(); - Ok(serde_json::json!({ "success": true, "count": repos.len(), "repos": repos })) - } - - fn get_code_metrics(&self, repo_id: &str) -> anyhow::Result { - let conn = self.conn()?; - match crate::registry::metrics::get_code_metrics(&conn, repo_id)? { - Some(m) => Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "total_lines": m.total_lines, - "source_lines": m.source_lines, - "test_lines": m.test_lines, - "comment_lines": m.comment_lines, - "file_count": m.file_count, - "language_breakdown": m.language_breakdown, - "updated_at": m.updated_at.to_rfc3339() - })), - None => { - Ok(serde_json::json!({ "success": false, "error": "No metrics found for repo" })) - } - } - } - - fn get_health(&self, repo_id: &str) -> anyhow::Result { - let conn = self.conn()?; - match crate::registry::health::get_health(&conn, repo_id)? { - Some(h) => Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "status": h.status, - "ahead": h.ahead, - "behind": h.behind, - "checked_at": h.checked_at.to_rfc3339() - })), - None => Ok(serde_json::json!({ "success": false, "error": "No health data found" })), - } - } - - fn query_call_graph( - &self, - repo_id: &str, - callee: Option<&str>, - caller: Option<&str>, - file: Option<&str>, - limit: usize, - ) -> anyhow::Result { - let conn = self.conn()?; - let edges = crate::registry::call_graph::query_call_edges( - &conn, - repo_id, - callee.filter(|s| !s.is_empty()), - caller.filter(|s| !s.is_empty()), - file.filter(|s| !s.is_empty()), - limit, - )?; - let calls: Vec = edges - .into_iter() - .map(|e| { - serde_json::json!({ - "caller_file": e.caller_file, - "caller_symbol": e.caller_symbol, - "caller_line": e.caller_line, - "callee_name": e.callee_name, - }) - }) - .collect(); - Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "count": calls.len(), - "calls": calls - })) - } - - fn query_dependencies( - &self, - repo_id: &str, - direction: &str, - relation_type: Option<&str>, - ) -> anyhow::Result { - let conn = self.conn()?; - let rel_filter = relation_type.filter(|s| !s.is_empty()); - let label = if direction == "incoming" || direction == "reverse" { - "reverse dependencies" - } else { - "dependencies" - }; - let rows = if direction == "incoming" || direction == "reverse" { - crate::dependency_graph::list_reverse_dependencies(&conn, repo_id)? - } else { - crate::dependency_graph::list_dependencies(&conn, repo_id)? - }; - let deps: Vec = rows - .into_iter() - .filter(|(_, rel, _)| rel_filter.is_none_or(|f| f == rel)) - .map(|(id, rel, conf)| { - serde_json::json!({ - "repo_id": id, - "relation_type": rel, - "confidence": conf, - }) - }) - .collect(); - Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "direction": direction, - "label": label, - "count": deps.len(), - "dependencies": deps - })) - } - - fn query_code_symbols( - &self, - repo_id: &str, - name: Option<&str>, - symbol_type: Option<&str>, - file: Option<&str>, - limit: usize, - ) -> anyhow::Result { - let conn = self.conn()?; - let mut sql = String::from( - "SELECT file_path, symbol_type, name, line_start, line_end, signature \ - FROM code_symbols WHERE repo_id = ?1", - ); - let mut params: Vec> = vec![Box::new(repo_id.to_string())]; - if let Some(ty) = symbol_type.filter(|s| !s.is_empty()) { - sql.push_str(" AND symbol_type = ?"); - sql.push_str(&(params.len() + 1).to_string()); - params.push(Box::new(ty.to_string())); - } - if let Some(n) = name.filter(|s| !s.is_empty()) { - sql.push_str(" AND name LIKE ?"); - sql.push_str(&(params.len() + 1).to_string()); - params.push(Box::new(format!("%{}%", n))); - } - if let Some(f) = file.filter(|s| !s.is_empty()) { - sql.push_str(" AND file_path LIKE ?"); - sql.push_str(&(params.len() + 1).to_string()); - params.push(Box::new(format!("%{}%", f))); - } - sql.push_str(&format!(" ORDER BY file_path, line_start LIMIT {}", limit.min(200))); - - let param_refs: Vec<&dyn rusqlite::ToSql> = params.iter().map(|p| p.as_ref()).collect(); - let mut stmt = conn.prepare(&sql)?; - let rows = stmt.query_map(rusqlite::params_from_iter(param_refs), |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, String>(2)?, - row.get::<_, i64>(3)?, - row.get::<_, i64>(4)?, - row.get::<_, Option>(5)?, - )) - })?; - - let mut symbols = Vec::new(); - for row in rows { - symbols.push(row?); - } - - let out: Vec = symbols - .iter() - .map(|(fp, st, n, ls, le, sig)| { - serde_json::json!({ - "file_path": fp, - "symbol_type": st, - "name": n, - "line_start": ls, - "line_end": le, - "signature": sig, - }) - }) - .collect(); - Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "count": out.len(), - "symbols": out - })) - } - - fn query_dead_code( - &self, - repo_id: &str, - include_pub: bool, - limit: usize, - ) -> anyhow::Result { - let conn = self.conn()?; - let mut sql = String::from( - "SELECT file_path, name, line_start, signature \ - FROM code_symbols cs \ - WHERE cs.repo_id = ?1 AND cs.symbol_type = 'function' \ - AND NOT EXISTS ( \ - SELECT 1 FROM code_call_graph ccg \ - WHERE ccg.repo_id = cs.repo_id AND ccg.callee_name = cs.name \ - )", - ); - if !include_pub { - sql.push_str(" AND (cs.signature IS NULL OR cs.signature NOT LIKE 'pub%fn%')"); - } - sql.push_str(" AND cs.name != 'main'"); - // Exclude test functions — heuristic: name starts with 'test_' (Rust convention) - sql.push_str(" AND cs.name NOT LIKE 'test_%'"); - // Exclude functions in tests.rs files (Rust unit-test modules) - sql.push_str( - " AND cs.file_path NOT LIKE '%/tests.rs' AND cs.file_path NOT LIKE '%\\tests.rs'", - ); - // Exclude functions with #[test] or #[tokio::test] attributes (tree-sitter extracted) - sql.push_str(" AND (cs.attributes IS NULL OR cs.attributes NOT LIKE '%#[test]%')"); - sql.push_str(&format!(" ORDER BY cs.file_path, cs.line_start LIMIT {}", limit.min(200))); - - let mut stmt = conn.prepare(&sql)?; - let rows = stmt.query_map([repo_id], |row| { - Ok(( - row.get::<_, String>(0)?, - row.get::<_, String>(1)?, - row.get::<_, i64>(2)?, - row.get::<_, Option>(3)?, - )) - })?; - - let mut dead = Vec::new(); - for row in rows { - dead.push(row?); - } - - let out: Vec = dead - .iter() - .map(|(fp, n, line, sig)| { - serde_json::json!({ - "file_path": fp, - "name": n, - "line_start": line, - "signature": sig, - }) - }) - .collect(); - Ok(serde_json::json!({ - "success": true, - "repo_id": repo_id, - "count": out.len(), - "dead_functions": out - })) - } + Ok(RepairResult { + orphans: orphaned, + missing_from_index: missing, + }) } /// Test-only storage backend that uses an independent temporary directory. @@ -656,6 +343,11 @@ impl StorageBackend for TempStorageBackend { std::fs::create_dir_all(dir)?; Ok(dir.join("search_index")) } + fn symbol_index_path(&self) -> anyhow::Result { + let dir = self.dir.path(); + std::fs::create_dir_all(dir)?; + Ok(dir.join("symbol_index")) + } fn backup_dir(&self) -> anyhow::Result { Ok(self.dir.path().join("backups")) } @@ -705,8 +397,9 @@ mod tests { std::thread::sleep(std::time::Duration::from_millis(800)); // Repair should detect the orphan - let count = repair_tantivy_consistency_at(&index_path, &mut conn).unwrap(); - assert_eq!(count, 1); + let result = repair_tantivy_consistency_at(&index_path, &mut conn).unwrap(); + assert_eq!(result.orphans, 1); + assert_eq!(result.missing_from_index, 0); let orphan_exists: bool = conn .query_row("SELECT 1 FROM orphan_tantivy_docs WHERE repo_id = 'ghost_repo'", [], |_| { @@ -723,8 +416,9 @@ mod tests { ).unwrap(); // Repair should now find zero orphans and clear the record - let count2 = repair_tantivy_consistency_at(&index_path, &mut conn).unwrap(); - assert_eq!(count2, 0); + let result2 = repair_tantivy_consistency_at(&index_path, &mut conn).unwrap(); + assert_eq!(result2.orphans, 0); + assert_eq!(result2.missing_from_index, 0); let orphan_still_exists: bool = conn .query_row("SELECT 1 FROM orphan_tantivy_docs WHERE repo_id = 'ghost_repo'", [], |_| { diff --git a/src/sync.rs b/src/sync.rs index 5b0ec8b..01bfe7b 100644 --- a/src/sync.rs +++ b/src/sync.rs @@ -219,5 +219,17 @@ pub async fn run( Ok(()) } +impl crate::clients::SyncClient for crate::storage::AppContext { + async fn sync_repos( + &self, + dry_run: bool, + filter_tags: Option>, + ) -> anyhow::Result { + let conn = self.conn()?; + let filter_tags_str = filter_tags.as_deref().map(|v| v.join(",")); + crate::sync::run_json(&conn, dry_run, filter_tags_str.as_deref(), None, &self.i18n).await + } +} + #[cfg(test)] mod tests; diff --git a/src/test_utils.rs b/src/test_utils.rs index 1458437..d773176 100644 --- a/src/test_utils.rs +++ b/src/test_utils.rs @@ -5,8 +5,8 @@ use chrono::Utc; use std::path::PathBuf; /// Create an in-memory SQLite connection with the full devbase schema. -pub fn temp_db() -> rusqlite::Connection { - WorkspaceRegistry::init_in_memory().expect("failed to create in-memory db") +pub fn temp_db() -> anyhow::Result { + WorkspaceRegistry::init_in_memory() } /// Build a minimal RepoEntry fixture for tests. diff --git a/src/workflow/scheduler.rs b/src/workflow/scheduler.rs index b6c6ff3..0725422 100644 --- a/src/workflow/scheduler.rs +++ b/src/workflow/scheduler.rs @@ -15,8 +15,9 @@ pub fn build_schedule(wf: &WorkflowDefinition) -> anyhow::Result anyhow::Result = VecDeque::new(); for _ in 0..batch_size { - let id = queue.pop_front().expect("queue not empty: checked by while condition"); - let step = wf.steps.iter().find(|s| s.id == id).expect("step id must exist").clone(); + let id = queue + .pop_front() + .ok_or_else(|| anyhow::anyhow!("queue not empty: checked by while condition"))?; + let step = wf + .steps + .iter() + .find(|s| s.id == id) + .ok_or_else(|| anyhow::anyhow!("step id must exist"))? + .clone(); batch.push(step); processed += 1; if let Some(children) = adj.get(id) { for &child in children { - let deg = in_degree.get_mut(child).expect("child id initialized in in_degree"); + let deg = in_degree + .get_mut(child) + .ok_or_else(|| anyhow::anyhow!("child id initialized in in_degree"))?; *deg -= 1; if *deg == 0 { next_queue.push_back(child); diff --git a/tools/invariant-checks/run-checks.ps1 b/tools/invariant-checks/run-checks.ps1 new file mode 100644 index 0000000..1871156 --- /dev/null +++ b/tools/invariant-checks/run-checks.ps1 @@ -0,0 +1,242 @@ +#!/usr/bin/env pwsh +# SPDX-License-Identifier: MIT +# Copyright (c) 2026 juice094 +# devbase Architecture Invariant CI Checks +# Run from repo root: tools/invariant-checks/run-checks.ps1 + +$ErrorActionPreference = "Stop" +$script:Failed = 0 +$script:Passed = 0 +$script:Warnings = 0 + +function Write-CheckHeader($name) { + Write-Host "`n==> $name" -ForegroundColor Cyan +} + +function Report($status, $message) { + if ($status -eq "PASS") { + $script:Passed++ + Write-Host " [PASS] $message" -ForegroundColor Green + } elseif ($status -eq "WARN") { + $script:Warnings++ + Write-Host " [WARN] $message" -ForegroundColor Yellow + } else { + $script:Failed++ + Write-Host " [FAIL] $message" -ForegroundColor Red + } +} + +# --- Helper: compute line ranges of #[cfg(test)] blocks in a file --- +function Get-TestLineRanges($filePath) { + $ranges = @() + if (-not (Test-Path $filePath)) { return $ranges } + $lines = Get-Content $filePath + $inTest = $false + $testStart = -1 + $depth = 0 + for ($i = 0; $i -lt $lines.Count; $i++) { + $line = $lines[$i] + if (-not $inTest -and $line -match '^\s*#\[cfg\(test\)\]\s*$') { + $inTest = $true + $testStart = $i + $depth = 0 + continue + } + if ($inTest) { + $open = ([regex]::Matches($line, '\{')).Count + $close = ([regex]::Matches($line, '\}')).Count + $depth += $open - $close + if ($depth -le 0 -and ($i -gt $testStart)) { + $ranges += @{Start = $testStart; End = $i} + $inTest = $false + } + } + } + if ($inTest) { + $ranges += @{Start = $testStart; End = $lines.Count - 1} + } + return $ranges +} + +function Is-LineInTestRange($lineNum, $ranges) { + foreach ($r in $ranges) { + if ($lineNum -ge $r.Start -and $lineNum -le $r.End) { + return $true + } + } + return $false +} + +# --- G5: RF-6 — Detect NEW unwrap/expect/panic in production code --- +Write-CheckHeader "G5: RF-6 new unwrap/expect/panic detection" + +$diffFiles = cmd /c "git diff --name-only origin/main 2>nul" +if (-not $diffFiles) { + Report "PASS" "No changes since origin/main" +} else { + $newViolations = @() + foreach ($file in $diffFiles -split "`n") { + if ($file -notmatch '\.rs$') { continue } + if ($file -match 'tests?/|_test\.rs$|benches/|examples/') { continue } + + # Get test line ranges for the file + $testRanges = Get-TestLineRanges $file + + $diff = cmd /c "git diff -U0 origin/main -- `"$file`" 2>nul" + $lines = $diff -split "`n" + $currentLine = -1 + for ($i = 0; $i -lt $lines.Count; $i++) { + $line = $lines[$i] + # Parse hunk header to get base line number + if ($line -match '^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@') { + $currentLine = [int]$Matches[2] + continue + } + if ($line -match '^@@') { continue } + if ($line -match '^(diff|index|---|\+\+\+)') { $currentLine = -1; continue } + if ($currentLine -lt 0) { continue } + + if ($line.StartsWith('+') -and -not $line.StartsWith('+++')) { + $added = $line.Substring(1) + if ($added -match '^\s*//') { $currentLine++; continue } + + # Check if this line is inside a cfg(test) block + $lineNum = $currentLine - 1 # convert to 0-based index + if ($testRanges.Count -gt 0 -and (Is-LineInTestRange $lineNum $testRanges)) { + $currentLine++ + continue + } + + if ($added -match '(?