Design source of truth. This document is the original product/design brief for StorOps and should be kept in sync as the design evolves. Implementation lives under
rules/,src/storops/, andSKILL.mdat the repo root.
我们要开发一个面向 AI Coding Agent 的 Storage Operations Skill,项目名称暂定:
StorOps
Skill 名称:
storops
一句话定位:
Storage Operations for AI Agents.
目标不是重新实现一个 WizTree,也不是简单做一个"磁盘空间分析器"或"磁盘清理器"。
核心理念:
WizTree 负责"看见",StorOps 负责"理解、规划和行动"。
StorOps 应该让 Claude Code、Codex、OpenCode 等 Agent 能够安全地理解和管理本地存储空间。
目前已经存在一些相关项目,例如:
disk-space-analyzer-skillwiztree-mcpdisk-cleaner
这些项目已经解决了很多基础问题,因此不要重复造轮子。
尤其是:
- 不重新实现磁盘扫描器
- 不重新实现 WizTree
- 不重新实现 NTFS MFT 扫描
- 不把项目做成一个新的 GUI 磁盘分析软件
我们应该站在现有工具之上。
现有工具主要解决:
"哪些文件/目录占用了空间?"
StorOps 要进一步解决:
"这些空间是什么?为什么在那里?是否可以删除?是否应该迁移?迁移到哪里?怎么安全迁移?迁移后是否正常?"
因此 StorOps 的定位应该是:
Discover
↓
Understand
↓
Diagnose
↓
Recommend
↓
Plan
↓
Execute
↓
Verify
StorOps 不是传统 GUI 工具。所有能力都应该围绕 Agent 使用设计。
Agent 应该能够自然地处理:
"为什么我的 C 盘快满了?" "找出 C 盘最大的 20 个目录。" "哪些东西可以清理?" "我的 LM Studio 模型为什么占了这么多 C 盘空间?" "把可以迁移的 AI 模型迁移到 E 盘。" "帮我清理掉可以安全删除的缓存。"
默认情况下,所有操作都应该是只读的。例如:scan / inspect / search / identify / analyze / diagnose / recommend 都可以自动执行。
任何修改用户文件系统的操作:delete / move / rename / junction / symlink / configuration change 必须经过明确的用户确认。
Agent 不应该看到 C:\Users\xxx\.cache 就凭经验猜"这应该是 Hugging Face"。应该尽可能通过确定性规则、软件配置、环境变量、注册表或已知应用目录进行识别。
识别结果应该包含:application / purpose / category / owner / size / location / confidence / actionability,例如:
{
"path": "C:\\Users\\xxx\\.cache\\huggingface",
"application": "Hugging Face",
"category": "ai-model-cache",
"size": 54.2,
"unit": "GB",
"confidence": 0.98,
"deletable": true,
"migratable": true
}第一阶段主要针对 Windows,因为 WizTree 在 Windows / NTFS 环境下具有非常优秀的扫描性能。但架构不与 WizTree 强绑定。StorOps 将 WizTree 视为 Windows storage discovery backend,而不是整个 StorOps。
v2 更新:调度器已从 PowerShell(
scripts/lib/ScanBackend.psm1)迁移到 Python (src/storops/platform/base.py),下面描述的是当前实现;scripts/lib/ScanBackend.psm1及其兄弟模块已随 v2 发布删除。scripts/*.ps1兼容包装脚本本身也已在之后的版本里整体 移除——storopsCLI(python -m storops ...)现在是唯一的调用方式。详见 §25 及docs/plans/storops-v2-cross-platform-refactor.md。
StorOps 通过 src/storops/platform/base.py 里的一组工厂函数(get_scan_backend() 等)这一层
调度器,把"扫一个目录、拿到它的直属子项大小"这件事和具体用什么工具做完全解耦——这是全仓库
唯一做 sys.platform/platform.system() 判断的地方,core/ 和 cli.py 永远只拿到已经
选好的 backend 实例:
src/storops/platform/
base.py Protocol 定义(ScanBackend/CapacityProvider/CopyEngine/
LinkEngine)+ 调度器(get_scan_backend() 等工厂函数)
posix.py Linux/macOS 共用实现(容量/复制/链接机制在两个平台上
完全一致,故合并为一个模块,而非规划文档最初设想的
linux.py + macos.py 两个文件——真实差异只在
core/rules.py 的 token 展开表里,拆两份纯属过度抽象)
windows/ Windows 专属子包(WizTree/robocopy/Junction/ctypes 等
机制与 Linux/macOS 有实质差异,因此单独成包)
backends/ 具体扫描后端实现:wiztree.py / gdu.py / du.py
每个 backend 必须实现同一份契约(ScanBackend Protocol,方法签名和返回形状完全一致):
scan()返回一组标准化Entry对象:full_name, is_folder, size_bytes, allocated_bytes, modified, file_count, folder_count(allocated_bytes在 Linux/macOS backend 上目前等于size_bytes—— ext4/APFS 没有 NTFS MFT 那种统一暴露"逻辑大小 vs 实际占用块数"的简单途径, 这是一个已知的、可接受的精度取舍)。top_entries()(scan/inspect 的直接依赖)path_size()(cleanup-plan/migrate-plan 给单个已知路径称重)
调度逻辑(get_scan_backend()):
Windows -> WizTree(无 WizTree 时 fallback 到原生 os.scandir 遍历)
Linux/macOS + gdu -> Gdu
Linux/macOS 无 gdu -> Du(打印一次性能提示,但仍然可用)
core/ 里的所有编排逻辑(scan.py/cleanup.py/migrate.py)只依赖 platform 包导出的工厂
函数和 Protocol,从不直接 import 某个具体 backend —— 这样新增/更换一个平台的 backend 不需要碰
任何编排代码。core/rules.py/core/risk.py 完全不感知 backend,只消费上面这组标准化字段。
回退到 Du 时,只打印一条 warning 已知不够可靠——agent 能否捕获到 stderr 并不确定。所以每个
子命令的 --json 输出(以及 cleanup-plan/migrate-plan 落盘的 plan 文件)都额外带了 Backend
(当前选中的 backend 名)和 BackendAdvice(回退到 Du 时是一句建议装 gdu 的文本,否则是
null)两个字段 —— 这是结构化数据,agent 一定读得到,不用赌 stderr 有没有被带回来。SKILL.md
第 13 条要求 agent 在 BackendAdvice 非空时提醒用户一次,而不是每条命令都念叨。
du 是逐文件 stat() 遍历,单线程,瓶颈是 I/O 延迟而不是吞吐 —— 在 SSD/NVMe 上尤其浪费,因为
一次只发一个 syscall,队列深度打不满。WizTree 快是因为它绕过文件系统驱动直接读 NTFS MFT,这个技巧
在 ext4/APFS 上没有公开、稳定的等价物(ext4 可以用 debugfs 读裸块设备,但需要 root 且脆弱,
StorOps 不会这么做)。
能做到的、性价比最高的加速手段是并行遍历:gdu(Go,goroutine 并发扫描,内置 JSON 导出,单个跨平台静态二进制)是目前最接近"WizTree 替身"的选择 —— 检测优先级为:
$env:STOROPS_GDU_PATH(显式指定,呼应$env:STOROPS_WIZTREE_PATH的现有约定)- PATH 上的
gdu - 都没有 -> 回退到系统自带的
du(始终可用,但大目录树上明显更慢,打印一次警告提示安装 gdu)
du 分支需要同时兼容 GNU coreutils(--max-depth/-b)和 BSD/macOS(-d/-k)两套完全不同的
参数,backends/Du.psm1 在调用前探测 du --version 来决定用哪一套;两种情况都把深度限制原生传给
du 本身(而不是先全量扫描再在 PowerShell 里截断),避免"只要顶层几个目录的大小"却触发一次全盘遍历。
rules/windows.yaml(已有)、rules/linux.yaml、rules/macos.yaml 各自维护该平台"绝不允许自动
清理/迁移"的关键系统路径短路规则,Identify.psm1 始终把三个文件都加载 —— 不匹配当前平台的 token
(如 Linux 上出现 %SYSTEMROOT%)不会展开,规则自然不命中,不需要按平台条件加载。ai-models.yaml
/applications.yaml/caches.yaml 目前的 path_patterns 仍以 Windows token 为主;补齐 Linux/macOS
下同一批应用(LM Studio、Ollama、Docker、npm/pip 等)的路径是后续需要单独投入的工作量,不在这次
抽象层改动范围内。
不要操作 WizTree GUI,不要使用 GUI automation / mouse click / screenshot / OCR。直接使用 WizTree CLI。
基本流程:
StorOps → invoke WizTree CLI → export structured data → parse result → normalize → analyze
优先使用 WizTree 原生的 CLI / export 能力,充分利用其支持的 CSV export、file type information、percentage information、drive capacity、maximum depth、treemap/export capabilities。
同时注意控制导出数据量,不要在每次扫描时无脑导出整个磁盘的全部文件。优先:
drive summary → top directories → targeted drill-down → targeted search
减少:扫描时间 / CSV 大小 / 内存占用 / Agent context/token 消耗。
扫描指定磁盘或目录(如 scan C:、scan C:\Users、scan E:\AI)。输出:total capacity / used / free / top directories / largest files / file type distribution。
深入分析指定路径,支持逐层展开(如 inspect C:\Users\xxx\AppData\Local)。
支持:find files > 10GB、find *.gguf、find model files、find files older than 1 year、find directories named cache。
这是 StorOps 与现有磁盘分析工具的重要区别。尝试识别:
- AI / Development: LM Studio, Ollama, Hugging Face, ComfyUI, Docker, WSL, npm, pnpm, yarn, pip, uv, conda, Python, Visual Studio, JetBrains, VS Code, Git
- General Applications: Steam, Chrome, Edge, Discord, Adobe, etc.
识别结果应该尽可能告诉 Agent:What is it? Who owns it? Why does it exist? Can it be deleted? Can it be moved? How should it be moved? What happens if it is deleted?
现代 AI 开发环境非常容易产生大量存储占用,例如:LM Studio models / Hugging Face cache / Ollama models / ComfyUI models / Stable Diffusion models / PyTorch cache / CUDA cache / npm cache / pip cache / uv cache / Docker images / WSL VHDX。
不能简单把它们全部分类成 cache = safe delete,而应该区分:Delete / Move / Keep / Re-download required / Currently in use / Configuration required。
例如:
Hugging Face cache — 54 GB
Delete: Yes
Consequence: Models may need to be downloaded again.
Migration: Recommended.
Target: E:\AI\HuggingFace
用户可能会说"C 盘的 LM Studio 模型太大了,帮我迁到 E 盘"。StorOps 应该能够:
1. Identify application
2. Identify storage directory
3. Determine whether application is running
4. Recommend migration method
5. Ask user for confirmation
6. Stop application if necessary
7. Move data
8. Update application configuration
9. Verify data
10. Remove old data only after verification
对不支持修改存储路径的软件,可以考虑用 Junction(Windows 优先 Junction 而非 symbolic link)把旧路径指向新位置。必须:确认目标路径、确认源路径、确认数据已经完整迁移、确认应用没有运行、创建后进行验证。
- LOW:temporary files、known disposable logs、safe application cache
- MEDIUM:Hugging Face cache、npm/pip cache、browser cache、Docker unused layers(需要明确告诉用户删除后的后果)
- HIGH:application data、development environments、large model files、WSL virtual disks
- CRITICAL:Windows、System32、Program Files、unknown system files、user documents(默认禁止 Agent 自动删除)
不要直接执行删除,而应该先生成 Cleanup Plan,列出每一项的 size / risk / consequence / action,汇总 total reclaimable,只有得到用户明确确认之后才能执行。
任何写操作都必须支持验证,例如迁移后核对 file count / total size / expected files / target accessible / source no longer contains original data / junction works。验证失败时不要自动删除原始数据。
允许保存 snapshot,Agent 可以回答"为什么我这个月 C 盘少了 200GB"之类的问题,通过对比 snapshot 得出各类别的增量。
- 第一阶段:
scan_drive,inspect_path,find_large_files,search_files,extension_summary,identify_path,analyze_storage - 第二阶段:
recommend_cleanup,recommend_migration,generate_action_plan - 第三阶段:
move_path,create_junction,delete_path,update_configuration,verify_operation - 第四阶段:
create_snapshot,compare_snapshots,storage_growth
- Read(自动执行):scan, inspect, search, identify, analyze
- Plan(自动执行,但不能修改文件):recommend_cleanup, recommend_migration, generate_action_plan
- Write(必须用户确认):move, delete, rename, junction, configuration
SKILL.md 不应该只是说明"调用 WizTree",而应该定义 Agent 的行为规范,例如:
When user asks why disk space is low:
1. Scan the relevant drive.
2. Find top-level consumers.
3. Drill down into unusually large directories.
4. Identify known applications/caches/models.
5. Classify each result.
6. Recommend actions.
7. Never delete automatically.
When user asks to clean:
1. Analyze first.
2. Generate cleanup plan.
3. Explain consequences.
4. Ask for confirmation.
5. Execute only approved actions.
6. Verify.
Skill 应该尽量让 Agent 主动使用 StorOps,而不是让用户必须知道具体工具名称。
这是 v1(PowerShell)时代的原始结构,作为历史记录保留。v2 之后的当前目录结构见 §25 及
docs/plans/storops-v2-cross-platform-refactor.md§2.2——scripts/目录(含其兼容 包装脚本)已不存在,实现全部在src/storops/下。
storops/ (repo root)
├── SKILL.md
├── README.md
├── LICENSE
├── docs/
│ └── DESIGN.md
├── scripts/
│ ├── lib/
│ │ ├── ScanBackend.psm1 调度器,见 §4a
│ │ ├── backends/
│ │ │ ├── WizTree.psm1 Windows
│ │ │ ├── Gdu.psm1 Linux/macOS 首选
│ │ │ └── Du.psm1 Linux/macOS 兜底
│ │ ├── Common.psm1
│ │ ├── Identify.psm1
│ │ └── Risk.psm1
│ ├── scan.ps1
│ ├── inspect.ps1
│ ├── search.ps1
│ ├── identify.ps1
│ ├── cleanup-plan.ps1
│ ├── cleanup-execute.ps1
│ ├── migrate-plan.ps1
│ ├── migrate-execute.ps1
│ └── verify.ps1
├── rules/
│ ├── applications.yaml
│ ├── caches.yaml
│ ├── ai-models.yaml
│ ├── windows.yaml
│ ├── linux.yaml
│ └── macos.yaml
└── tests/
不要为了架构完整而过早复杂化;MCP server 属于后续阶段。跨平台 scan backend
(§4a)已经在这次改动里做了,但 ai-models.yaml/applications.yaml/
caches.yaml 的 Linux/macOS 路径覆盖仍是后续工作(§4c)。
StorOps 不是 disk-space-analyzer-skill 的 clone,也不是 wiztree-mcp 的 fork,也不是 disk-cleaner 的 clone。应该复用它们已经验证过的思路(WizTree CLI、CSV export、targeted scanning、structured analysis、risk classification),然后把价值集中在:Application Identification / AI Model Awareness / Migration / Action Planning / Verification / Agent-native Workflow。
- A. WizTree integration:find WizTree、invoke CLI、export structured data、parse data
- B. Disk analysis:scan drive、top directories、largest files、file extensions、drill-down
- C. Application identification:至少支持 LM Studio, Ollama, Hugging Face, ComfyUI, Docker, WSL, npm, pnpm, pip, uv
- D. Recommendations:KEEP / DELETE / MOVE / CHECK,并说明 risk / reason / consequence / recommended destination
- E. Safety:所有修改操作 confirmation required
- F. Verification:迁移和清理完成后必须验证
GUI、自己实现磁盘扫描器/MFT scanner、自动后台监控、自动定时清理、自动删除未知文件、复杂数据库、云端服务。
("跨平台完整支持"曾经也在这份名单里——§4a/§4b/§4c 记录的 scan-backend 抽象已经把
Linux/macOS 的核心扫描能力做出来了,是一次主动的范围扩展,而不是踩了这条非目标。
但"完整"两个字仍未达到:ai-models.yaml/applications.yaml/caches.yaml 的
Linux/macOS 应用规则、以及对应的 smoke test 覆盖,仍然是待办,见 §4c。)
StorOps scanning C:...
C: 930 GB Used: 891 GB Free: 39 GB
Largest consumers:
LM Studio models 87 GB
Hugging Face cache 54 GB
WSL VHDX 48 GB
Docker 31 GB
Windows 28 GB
Downloads 21 GB
LM Studio models — 87 GB
Identified: LM Studio
Recommended: MOVE → E:\AI\LMStudio\Models
Risk: LOW
Reason: Model files are large and portable.
Migration Plan
Source: C:\Users\xxx\.lmstudio\models Size: 87.2 GB
Target: E:\AI\LMStudio\models
Method: Application-supported path change
Steps:
1. Close LM Studio
2. Move models
3. Update model directory
4. Verify models
5. Remove old files
Proceed?
StorOps found 37.8 GB of low-risk cleanup candidates.
LOW RISK
Temp files 8.4 GB
npm cache 4.1 GB
pip cache 2.8 GB
MEDIUM RISK
Hugging Face cache 22.5 GB
Consequence: Models may need to be downloaded again.
I will only clean LOW RISK items unless you approve the Hugging Face cache separately.
Proceed with 15.3 GB cleanup?
- 分析优先,执行其次。
- 默认只读。
- 不要猜测文件用途。
- 不要因为名字叫 cache 就认为可以删除。
- 不要删除未知文件。
- 不要自动修改系统目录。
- 任何破坏性操作必须获得明确确认。
- 迁移完成后必须验证。
- 如果应用正在运行,不要直接移动其数据。
- 对于可能重新下载的大型 AI 模型,必须明确告知用户后果。
- 优先迁移,而不是删除用户有价值的数据。
- 尽可能使用应用官方支持的路径配置,而不是强制使用 Junction。
- 只有在应用不支持路径配置时,才考虑 Junction。
- 不要让 WizTree 成为整个架构的强依赖。
不是"我们能扫描 C 盘",而是:用户问"为什么 C 盘满了",Agent 可以从扫描结果一路追踪到具体的软件/缓存/模型,并给出可靠、可执行、安全的解决方案,整个过程都由 Agent 驱动:
C:\... → 87 GB → LM Studio → AI model storage → migratable → E:\AI\LMStudio
→ migration plan → user confirmation → move → verify
- 项目名称:StorOps
- Skill:
storops - Tagline:Storage Operations for AI Agents.
- 核心理念:See where your storage goes. Understand why. Move what matters. Clean what doesn't.
不要把它定位成 Disk Cleaner 或 Disk Analyzer,而应该定位成 Storage Operations layer for AI Agents:
StorOps
│
┌────────────┼────────────┐
│ │ │
Discover Understand Diagnose
│ │ │
WizTree Identify Analyze
│ │ │
└────────────┼────────────┘
│
Plan
│
┌─────────┴─────────┐
│ │
Migrate Clean
│ │
└─────────┬─────────┘
│
Verify
第一阶段重点不是"做更多功能",而是把这条 Agent workflow 做正确:优先复用 WizTree,把开发精力放在识别、智能判断、迁移规划、安全执行和验证上。
从某个版本起,StorOps 的实现语言从 PowerShell 迁移到了 Python(src/storops/),并获得了一个
统一的 storops CLI(storops scan/inspect/search/identify/cleanup/migrate/verify)取代原先
9 个互相独立的 .ps1 入口。这次重构的完整审计、架构方案、决策记录(依赖策略、兼容策略、
Linux/macOS 规则补齐方案等)都记录在
docs/plans/storops-v2-cross-platform-refactor.md——
本文档不重复那些细节,只记两条对理解现状最关键的结论:
- 本文档描述的产品设计/安全模型/能力边界本身没有变(三层安全模型、
Assert-not-critical兜底、未识别路径默认拒绝、迁移的 copy-verify-remove-relink 状态机等, §3/§8-§12 全部原样保留,只是实现语言换了);本节之前各处提到具体 PowerShell 模块 (Identify.psm1/Risk.psm1/ScanBackend.psm1等)的地方,应理解为对应逻辑现在位于src/storops/core/、src/storops/platform/下的同名 Python 模块(§4a 已更新为反映这一点)。 scripts/*.ps1曾经在 v2 发布时作为强制交付项与新 CLI 同版本一起提供(薄包装脚本,把 参数翻译成storopsCLI flag 后转调python -m storops),但按规划文档 §2.10 里"存在到用户 明确决定不再需要为止"的约定,兼容包装层已在之后被整体移除——storopsCLI(python -m storops ...)现在是唯一的调用方式。详见该规划文档 §2.10。
可执行文件:WizTree64.exe(或 WizTree.exe,32 位)。命令行导出用法:
WizTree64.exe "<drive-or-folder>" /export="<output.csv>" [options]
关键参数:
| 参数 | 说明 |
|---|---|
| `/admin=0 | 1` |
| `/exportfolders=0 | 1` |
| `/exportfiles=0 | 1` |
/exportmaxdepth=n |
限制导出的目录深度,0 为不限制 |
/sortby=n |
0=name, 1=size desc, 2=allocated desc, 3=date desc |
/filter="spec" |
只包含匹配的文件(如 *.gguf) |
/filterexclude="spec" |
排除匹配的文件 |
| `/filterfullpath=0 | 1` |
/exportallsizes=1 |
导出目录“自身”大小(不含子目录) |
/exportpercentofparent=1 |
导出相对父目录的百分比 |
/exportdrivecapacity=1 |
导出盘符总容量 |
CSV 导出列:File Name, Size, Allocated, Modified, Attributes, Files, Folders。目录名以 \ 结尾;Size/Allocated 对目录是递归总和;Attributes 为位掩码(1=只读, 2=隐藏, 4=系统, 32=归档, 2048=压缩)。
StorOps 通过组合 /exportfolders=1 /exportfiles=0 /exportmaxdepth=N /sortby=1 来实现"drive summary → top directories → targeted drill-down"的分层、受控扫描,避免一次性导出整盘全部文件。