From 81507377afe9640650dbbec43bd91419ed760d9a Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 01/15] docs(site): add the repository README and the design-system reference MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:docs(site): add the repository README and the design-system reference --- README.md | 363 +++++++++++++----------------------------------------- 1 file changed, 83 insertions(+), 280 deletions(-) diff --git a/README.md b/README.md index b529f9c3..ad46d853 100644 --- a/README.md +++ b/README.md @@ -4,336 +4,139 @@
-# 🧬 Distilly +# Distilly -**Formerly: Colleague Skill / colleague-skill.** +### Distill how they think into Person Profiles for Agents. -### Distill a person's experience, judgment, voice, and ways of working into a reusable Person Profile for AI agents and compatible bots. - -**Messages · documents · interviews · public sources → Distilly → Person Profile → Agent / Bot** +**Colleague Skill / colleague-skill (original name)** [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE) -[![Python 3.9+](https://img.shields.io/badge/Python-3.9%2B-blue.svg)](https://python.org) -[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io) -[![Stars](https://img.shields.io/github/stars/titanwings/colleague-skill?style=social)](https://github.com/titanwings/distilly/stargazers) - -[![Discord](https://img.shields.io/badge/Discord-Join%20Community-5865F2?logo=discord&logoColor=white)](https://discord.gg/NVX66RxWZv) - -
- - - -
- -🧑‍💼  Your colleague quit, your mentor graduated, your teammate transferred — taking their whole playbook and context with them?
-💞  Your family, old friends, partner drifting apart — and you want to hold on to the way it felt to be with them?
-🌟  Your favorite author, idol, thinker you'll never meet — but you want to know what they'd say about your question? - -
- -### ✨ One project, many kinds of people. - -
- -Distilly is the person-modeling layer for agents. It turns the materials you provide into a portable, source-grounded Person Profile built from observable experience, decision patterns, expression, and ways of working; it does not claim to clone the person behind them. - -Colleagues · partners · family · old friends · idols · public figures · fictional characters — even yourself - -**Source material + your description → a source-grounded Person Profile → your Agent or compatible Bot** - -> A Person Profile is the reusable output. The current release packages each profile as an Agent Skill so supported hosts can install and invoke it. The canonical creator Skill is named `distilly`; install it in a `distilly` directory. The former name above remains for search continuity and project history. - -
- -[🆕 What Distilly does](#-what-distilly-does-today) · [📦 Data Sources](#-supported-data-sources) · [⚡ Install](#-install) · [🚀 Usage](#-usage) · [✨ Demo](#-demo) · [📝 Citation](#-citation) · [💬 Discord](https://discord.gg/NVX66RxWZv) - -[**Chinese**](docs/lang/README_ZH.md) · [**Spanish**](docs/lang/README_ES.md) · [**German**](docs/lang/README_DE.md) · [**Japanese**](docs/lang/README_JA.md) · [**Russian**](docs/lang/README_RU.md) · [**Portuguese**](docs/lang/README_PT.md) · [**Korean**](docs/lang/README_KO.md) - - - ---- - -
- -### 🎉 2026.08.13 Milestone — **the project has passed 20K ⭐!** - -Massive thanks to everyone who starred — we'll keep shipping, keep distilling. - -
- -> 🧬 **2026.08.24 Update** — The creator is now named **Distilly** end to end and documents native local Skill discovery for Claude Code, Hermes, OpenClaw, Codex, DeepSeek Harness, Pi, Grok Build, and OpenCode. Grok Bot is listed separately as a saved-Skill workflow preview. - -> 📝 **2026.06.01 Update** — **[The COLLEAGUE.SKILL technical report](https://arxiv.org/pdf/2605.31264) is now available**. The most rewarding part was not simply publishing a paper, but seeing the community grow the gallery to 215 skills contributed by 165 people, with more than 100,000 stars across the skill cards. The paper's Acknowledgements explicitly recognize every community contributor. - -> 🗺️ **2026.04.13** — **The Distilly Roadmap is live!** What began as Colleague Skill is growing beyond colleagues: distill people into Skills that Agents can reuse. 👉 **[Full Roadmap](ROADMAP.md)** · **[💬 Discord](https://discord.gg/NVX66RxWZv)** - -> 🌐 **2026.04.07** — Community gallery is live! Any skill / meta-skill can drive traffic directly to your own GitHub repo. No middleman. 👉 **[titanwings.github.io/colleague-skill-site](https://titanwings.github.io/colleague-skill-site/)** - -
- -Created by [@titanwings](https://github.com/titanwings) +[![Node.js](https://img.shields.io/badge/Node.js-22.19%2B-339933.svg)](https://nodejs.org/) +[![Codex Preview](https://img.shields.io/badge/Codex-Developer%20Preview-black)](https://github.com/titanwings/distilly/tree/distilly-plugin)
---- - -## 🆕 What Distilly does today - -### 1️⃣ From Colleague Skill to Distilly - -The project is no longer limited to the colleague scenario. Its `distilly` creator builds source-grounded Person Profiles for three person families with one workflow, then packages each profile as an Agent Skill. - -### 2️⃣ Three character families - - - - - - - - - - - - - - - - - - - - - -
🧑‍💼 colleague💞 relationship🌟 celebrity
Coworkers · mentors · teammates · up/downstream partnersExes · partners · parents · friends · close familyPublic figures · creators · public voices · fictional characters
Builds a Work Skill + Persona from material-derived technical standards, workflows, expression, and workplace behavior. Supports Lark / DingTalk / Slack collection.Organizes material-derived expression patterns, emotional triggers, conflict patterns, and repair patterns into a reusable Persona Skill.Ships with a six-dimension research toolchain (subtitles → transcript cleanup → research merge → quality check) for organizing observable decisions, expression, and mental models.
- -Each family has its own source-collection strategy, analysis dimensions, and Person Profile structure. - -### 3️⃣ More Agent hosts - -The old version only ran in Claude Code. Distilly now supports native local Skill discovery across eight agent hosts. - - - - - - - - - - - - - - -
Claude CodeHermes AgentOpenClawCodex
DeepSeek HarnessPi coding agentGrok BuildOpenCode
- -**Grok Bot preview:** Grok Bot supports saved/private Skills, but its official docs do not describe direct local `SKILL.md` imports. Distilly's workflow can be migrated manually into a saved Skill; direct repo installation is not yet verified. +Distilly is a local-first product for turning a person's source material, working habits, judgment, and voice into a versioned **Person Profile for Agents**. The profile can be recalled temporarily during a run or explicitly installed as a long-lived host Skill. The storage authority stays local; no additional model API key is required. -Each generated Person Profile is packaged as an Agent Skill and can be installed into any supported host. +This `distilly-plugin` branch carries the unreleased `0.1.0-preview.1` Developer Preview; the repository's default branch is `dot-skill`, the separate legacy implementation, so a bare clone lands on that line instead of this one. Codex, OpenClaw `2026.3.24`, and Hermes `v0.9.0` each have an immutable real-host transport-capacity fixture. The OpenClaw and Hermes measurements use a deterministic synthetic fixture server through the real host executable, model, and MCP transport; they do not by themselves certify packaged restart or the full product lifecycle. Setup remains fail-closed for any unrecorded host version or changed release tuple. This branch is not a tagged release or an npm package yet. ---- +[Chinese](docs/lang/README_ZH.md) · [Español](docs/lang/README_ES.md) · [Deutsch](docs/lang/README_DE.md) · [日本語](docs/lang/README_JA.md) · [한국어](docs/lang/README_KO.md) · [Português](docs/lang/README_PT.md) · [Русский](docs/lang/README_RU.md) -## 📦 Supported Data Sources +## Install the Developer Preview -| Logo | Source | Messages | Docs / Wiki | Notes | -|:----:|--------|:--------:|:-----------:|-------| -| Lark | Lark (auto) | ✅ API | ✅ | Just enter a name, fully automatic | -| DingTalk | DingTalk (auto) | ⚠️ Browser | ✅ | DingTalk API doesn't support message history | -| Slack | Slack (auto) | ✅ API | — | Requires admin to install Bot; free plan limited to 90 days | -| X | Public X posts | ✅ API | — | Optional, bounded celebrity research candidates through metered third-party service Xquik | -| WeChat | WeChat chat history | ✅ SQLite | — | Export first with WeChatMsg or PyWxDump | -| 📄 | PDF / Images / Screenshots | — | ✅ | Manual upload | -| Lark | Lark JSON export | ✅ | ✅ | Manual upload | -| ✉️ | Email `.eml` / `.mbox` | ✅ | — | Manual upload | -| 📝 | Markdown / direct paste | ✅ | ✅ | Manual input | +### For an agent ---- +Give your coding agent the following task and let it run the commands in a fresh checkout: -## ⚡ Install +> Install the Distilly Developer Preview from the `distilly-plugin` branch, build it with Node 22.19+ (or Node 24), run `distilly setup --host codex`, run `distilly doctor --host codex`, and report the result. Do not modify another branch. -### 🤖 For Agents +The exact checkout and setup commands are shown below so the agent can verify every step. -Open any supported local Agent host and send: +### For a human -> Install Distilly from `https://github.com/titanwings/distilly`, then verify that this host can discover it. - -The Agent installs Distilly as a Skill named `distilly` in the correct host directory. - -### 👤 For Humans - -Clone Distilly into the Skills directory used by your host: +Requirements: Node.js `22.19+` or `24`, pnpm `10.32+`, and a locally installed Codex CLI. From a terminal: ```bash -git clone https://github.com/titanwings/distilly +git clone --branch distilly-plugin https://github.com/titanwings/distilly.git +cd distilly +corepack enable +pnpm install --frozen-lockfile +pnpm run build +node packages/cli/lib/bin.js setup --host codex +node packages/cli/lib/bin.js doctor --host codex ``` -Host paths, migration, Windows, generated-profile installation, and credential setup are in the **[Install Guide](INSTALL_EN.md)**. - ---- - -## 🚀 Usage - -In your Agent, say: - -> Use Distilly to create a Person Profile for ``. - -Then: - -1. Choose `colleague`, `relationship`, or `celebrity`. -2. Add a description or source material. Every field is optional. -3. Review the result and let Distilly generate the profile. - -The result is packaged as an Agent Skill named `{character}-{slug}`. For the optional celebrity research pipeline and advanced tools, see the **[Install Guide](INSTALL_EN.md)**. - ---- +Restart Codex after setup. The launcher registers the self-contained Plugin and its five MCP tools. To remove the host integration while keeping all local people, profiles, and source data: -## ✨ Demo - -One from each family. - -
- -### 🌟 celebrity — distilling Andrej Karpathy - -> Six-dimension research: works · interviews · decisions · expression DNA · external evaluations · timeline
-> Case adapted from community work [alchaincyf/karpathy-skill](https://github.com/alchaincyf/karpathy-skill) - -``` -User ❯ Everyone's talking about Agents now — what do you think the 2026 AI Agent looks like? - -celebrity.skill ❯ Everyone's doing demos, nobody's doing eval. - The Agent bottleneck was never model intelligence — it's scaffolding: - a strict evaluation harness, fine-grained error replay, - controllable tool-calling budgets, clean failure-mode analysis. - Most products are still stuck at "wrap an API + a few tools + a system prompt." - That's a demo, not an agent. - In 2026 I'm not betting on the strongest model — I'm betting on the best - scaffolding and the strictest eval. +```bash +node packages/cli/lib/bin.js uninstall --host codex ``` -
- -### 🧑‍💼 colleague — distilling a ByteDance L2-1 backend - -> Input: `ByteDance L2-1 backend engineer, INTJ, blame-shifter, ByteDance-style` +To install one approved profile as a persistent Skill after a profile has been created, use its exact subject id: +```bash +node packages/cli/lib/bin.js install subject_<32 lowercase hex characters> --host codex ``` -User ❯ Can you review this API design? -colleague.skill ❯ Hold on — what's the impact? You haven't explained the context. - (after reading) N+1 query, fix it. Use the standard - {code, message, data} response format. That's the spec, - don't ask why. +## Host compatibility and explicit Legacy fallback -User ❯ This bug was introduced by you, right? +Codex uses the native Plugin preview above. The Preview also includes compatibility bindings for OpenClaw and Hermes: -colleague.skill ❯ Does the timeline match? That feature touched multiple places, - there were other changes too. -``` +- **OpenClaw** loads the Claude-compatible bundle from `~/.openclaw/extensions/distilly` and its real `.mcp.json`. Check discovery with `openclaw plugins inspect distilly --json`. +- **Hermes** installs the canonical Skill at `~/.hermes/skills/distilly`, a managed wrapper at `~/.distilly/bin/distilly-hermes`, and an MCP entry in `~/.hermes/config.yaml`. `resources` and `prompts` are disabled so the exposed surface remains five tools; check it with `hermes mcp test distilly`. -
+The CLI recognizes both hosts and enables setup when the installed version matches the recorded real-host transport fixture. The current net budgets, measured in isolated clean sessions with `openai-codex/gpt-5.4`, are 65,536 serialized bytes for OpenClaw and 49,752 for Hermes (the same conservative byte/token accounting used by the Codex fixture). These are transport/value lower bounds for the recorded probe, not a guarantee of remaining context in every model or user session. Any unrecorded version, release digest, tool descriptor, or serializer tuple returns `host_unsupported` before writing an unverified integration. There is no automatic switch to the legacy implementation. -### 💞 relationship — distilling someone you have a crush on +Until a host has a verified Plugin binding, you can explicitly choose the maintained `dot-skill` branch as a **Legacy Skill compatibility mode**: -> Upload half a year of chat logs + "sensitive, quiet but stubborn, will actually reply seriously when it matters" +> Install Distilly in Legacy Skill compatibility mode from the `dot-skill` branch into this host's normal Skills directory, using a clean checkout whose final directory is named `distilly`. Verify discovery and report the installed Git commit. Do not run Plugin setup or claim SQLite, five-tool MCP, Panel, or Plugin lifecycle support. -``` -User ❯ Did you think about me today? +For a manual install, replace `` with the complete final path in the [detailed install guide](INSTALL.md), including the last `distilly` component, and create its parent first: -relationship.skill ❯ ...I did, a little bit. Why are you asking? +```bash +git clone --single-branch --branch dot-skill --depth 1 \ + https://github.com/titanwings/distilly.git \ + +git -C rev-parse HEAD ``` -
- -📚 More real-world cases in the **[community gallery](https://titanwings.github.io/colleague-skill-site/)** — 100+ skills and counting - -
- ---- - -## 🔧 Features +This is an explicit, separate file-based implementation—not an automatic runtime fallback. It does not share a supported data model with the Plugin, and a failed Plugin preflight never switches modes. The compatibility promise currently covers local files and pasted text only. Do not enable legacy collectors while the Plugin uses the same home directory: current legacy collectors can write credential configuration into the same `~/.distilly/` namespace and remain outside the Preview's reviewed security boundary. Keep exactly one `distilly` active in any host discovery scope and verify which copy the host loaded. -### 🧱 Generated Skill Structure +## The first usable flow -Distilly's current creator uses **Persona** as the universal base, with family-specific modules layered on top: +On Codex, the complete flow below is verified. OpenClaw `2026.3.24` and Hermes `v0.9.0` have the same briefing transport path verified against their recorded capacity fixture; their packaged restart, long-lived Skill, and uninstall lifecycle checks remain separate. Restart the selected host and ask it to research and distill a person. Supply only the files, text, or public URLs you want included. Distilly then: -| Family | Persona Content | Additional Modules | -|--------|-----------------|-------------------| -| 🧑‍💼 **colleague** | 6-layer personality: hard rules → identity → expression → decisions → interpersonal → Correction | ➕ **Work Skill**: scope, workflow, output preferences, experience knowledge base | -| 💞 **relationship** | Expression DNA · emotional triggers · conflict pattern · repair pattern | — | -| 🌟 **celebrity** | Mental models · decision heuristics · expression DNA · external-evaluation contrast | ➕ Six-dimension research dossier (works / interviews / decisions / timeline...) | +1. resolves or creates the person; +2. imports the selected material with deterministic local parsers; +3. creates a pending research job and a complete evidence-bound briefing; +4. commits a versioned Person Profile; +5. returns the profile or a complete temporary prompt for the current run; +6. accepts an explicit correction and sends a candidate to review; +7. lets you promote, reject, or roll back the candidate in the local Panel; and +8. installs the approved profile as a self-contained host Skill when you ask it to. -> **Execution**: Receive task → Persona selects material-derived preferences and tone → Additional modules fill in execution detail → Produce a source-grounded response +The model-facing surface remains exactly five MCP tools: -### 🧬 Evolution +`distilly_get` · `distilly_ingest` · `distilly_pending` · `distilly_commit` · `distilly_correct` -- 📥 **Append files** → auto-analyze delta → merge into relevant sections, never overwrite existing conclusions -- 💬 **Conversation correction** → say "they wouldn't do that, they'd be xxx" → writes to the Correction layer, takes effect immediately -- 🕰️ **Version control** → auto-archive on every update, rollback to any previous version -- 🔬 **Celebrity research pipeline** → subtitles → transcript cleanup → six-dimension research → quality check +Distilly never silently truncates a complete briefing or profile prompt. If a verified host budget cannot carry the complete value, it reports a bounded capacity error with measurements and keeps the stored data unchanged. ---- +## Host status -## ⚠️ Notes +| Host | Native Plugin | Current compatibility route | +| --- | --- | --- | +| Codex | Fully verified in this release branch | Native Plugin | +| Claude Code | Binding included; exact host fixture still needed | Explicit `dot-skill` Legacy Skill | +| OpenClaw | Transport-capacity fixture recorded for `2026.3.24` (65,536-byte net budget); lifecycle pending | Claude-compatible bundle + discovery smoke | +| Hermes | Transport-capacity fixture recorded for `v0.9.0` (49,752-byte net budget); lifecycle pending | Managed Skill + MCP configuration | +| DeepSeek Harness (DSH) | Community binding planned | Explicit `dot-skill` Legacy Skill | +| Pi agent | Community binding planned | Explicit `dot-skill` Legacy Skill | +| Grok Build | Community binding planned | Explicit `dot-skill` Legacy Skill | +| OpenCode | Community binding planned | Explicit `dot-skill` Legacy Skill | +| Grok Bot | Community binding planned | Manual saved/private Skill only; local repository import is not claimed | -**Source material quality = Person Profile quality** — and quality sources differ across families: +Host compatibility is a binding concern. Legacy Skill discovery is useful continuity, but it does not make a host a verified Plugin target. -| Family | Source priority (high → low) | -|--------|------------------------------| -| 🧑‍💼 **colleague** | Their **own long-form writing** (design docs / review comments) **›** **decision-making replies** **›** casual group chat | -| 💞 **relationship** | Complete chat history **›** letters / social posts / diaries **›** third-party descriptions | -| 🌟 **celebrity** | First-person books / blogs / long interviews **›** decision records (launches, commits, Q&A) **›** verified first-person short-form posts **›** third-party commentary | +## Local material formats -- **colleague** Lark-compatible auto-collection: requires adding the App bot to relevant group chats -- **relationship**: longer time spans are better; material covering both conflict and repair is ideal -- **celebrity**: avoid feeding only second-hand interpretations -- This is still a demo version — please file issues if you find bugs! +The first Preview accepts explicit local `TXT`, `Markdown`, `JSON`, and `SRT/VTT` files. It also accepts pasted text and public URLs through the host's visible research flow. Files are read only from the paths or sources the user supplies; symlinked selected files and duplicate file names are rejected. PDF, email, provider exports, and hosted connectors are follow-up work. ---- +## 📣 2026-09 update: help expand coding-agent Plugins -## 📄 Technical Report +Codex, OpenClaw, and Hermes now have real host/version capacity fixtures. We need community support to provide the same evidence for **Claude Code, DeepSeek Harness (DSH), Pi agent, Grok Build, OpenCode, and Grok Bot**, then to build and validate their coding-agent Plugin packages. I will actively review those contributions and keep the public contracts, release digests, and host behavior aligned. -> **[COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation](https://arxiv.org/pdf/2605.31264)** ([arXiv](https://arxiv.org/abs/2605.31264) · [arXiv PDF](https://arxiv.org/pdf/2605.31264)) -> -> This is the paper for **COLLEAGUE.SKILL / colleague-skill**, Distilly's predecessor. It covers the Work Skill + Persona two-layer architecture, multi-source data collection, and Skill generation mechanics — the theoretical foundation for today's `colleague` family. Separate papers on the relationship / celebrity family extensions are planned. +See the full call for contributors in [UPDATES.md](UPDATES.md) and the current priorities in [ROADMAP.md](ROADMAP.md). ---- +## Project documents -## 📝 Citation +- [Detailed Preview installation](INSTALL.md) +- [Changelog](CHANGELOG.md) +- [Architecture and shipped-state map](docs/architecture.md) +- [Testing contract](docs/testing.md) +- [Development workflow](docs/development.md) +- [Design corpus](docs/design/README.md) +- [Release manifest](plugins/release-manifest.json) +- [Contributing](CONTRIBUTING.md) -If you use **Distilly** or **COLLEAGUE.SKILL** in your research or applications, please cite the technical report: +Distilly is released under the [MIT License](LICENSE). Created by [@titanwings](https://github.com/titanwings). -```bibtex -@misc{zhou2026colleagueskill, - title = {COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation}, - author = {Tianyi Zhou and Dongrui Liu and Leitao Yuan and Jing Shao and Xia Hu}, - year = {2026}, - eprint = {2605.31264}, - archivePrefix = {arXiv}, - primaryClass = {cs.AI}, - url = {https://arxiv.org/abs/2605.31264} -} -``` - -You can also use the machine-readable citation metadata in [CITATION.cff](CITATION.cff). - ---- - -## ⭐ Star History - - - - - - Star History Chart - - - ---- - -
- -**MIT License** © [titanwings](https://github.com/titanwings) - -
From 66e409564f0c7204628467712a3296c76bb6ccd9 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 02/15] docs(v2): freeze the dot-skill v2 contract before parallel work MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:docs(v2): freeze the dot-skill v2 contract before parallel work --- .gitignore | 15 +++++++++++++++ docs/v2/CONTRACT.md | 1 + 2 files changed, 16 insertions(+) diff --git a/.gitignore b/.gitignore index 08b33f67..71290c97 100644 --- a/.gitignore +++ b/.gitignore @@ -29,7 +29,19 @@ playwright-data/ /tmp/feishu_*.txt /tmp/email_*.txt /tmp/dingtalk_*.txt + +# Every `knowledge/` is user data, wherever the person directory lives. This is an +# unanchored directory pattern, so it also matches source and fixture directories of +# the same name; the re-includes below must come *after* it, otherwise the later +# `knowledge/` rule excludes them again. Tracked files are unaffected either way, +# so only a newly added module would be lost — and it would be lost silently. knowledge/ +# ...except the synthetic ledger fixtures, which are test data and must be tracked +!src/derive/fixtures/**/knowledge/ +!src/derive/fixtures/**/knowledge/** +# ...and except the source modules: `src/knowledge/**` is code, not an export +!src/knowledge/ +!src/knowledge/** # OS .DS_Store @@ -37,4 +49,7 @@ Thumbs.db # 本地证据(截图/回执/diff),不入库 dst-evidence/ +# 根级渲染产物:契约 §2 把页面放在 evidence/renders/。 +# 前导斜杠是必须的 —— 裸写 evidence/ 会匹配 docs/evidence/,那里的文字证据是要入库的。 +/evidence/ *.evidence.local diff --git a/docs/v2/CONTRACT.md b/docs/v2/CONTRACT.md index 19db858d..f9e175bc 100644 --- a/docs/v2/CONTRACT.md +++ b/docs/v2/CONTRACT.md @@ -85,3 +85,4 @@ skills/// ## 6. 双语 prompt 与用户可见文档:**单文件双语**,中文段 → `---` → `## English`。 + From 9bbc9598a9cd07eb727bec7a9f2ef19d9e1b3025 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 03/15] docs(v2): note that the host matrix ships from the integration branch MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:docs(v2): note that the host matrix ships from the integration branch --- docs/v2/STATUS.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/v2/STATUS.md b/docs/v2/STATUS.md index c90cf2d9..7e95bfbc 100644 --- a/docs/v2/STATUS.md +++ b/docs/v2/STATUS.md @@ -9,12 +9,12 @@ | 3 | `ds/03-render` | 模板碎片 + 生成物 + `--check` 防漂移;`view check/render`;`visual-check` 八项 | — | `node scripts/generate-template.mjs --check`;`node scripts/visual-check.mjs ` | | 4 | `ds/04-prompts` | `SKILL.md` 五步 + 每个 prompt"必须/禁止/回执" + 三个新 prompt + `prompt-lint` | — | `node scripts/prompt-lint.mjs`;`node --test tests/prompt-contract.test.mjs` | | 5 | `ds/05-agents` | `docs/v2/HOSTS.md` 双语文档 + INSTALL/README 宿主章节 + 断言测试(矩阵本身已由维护者落在 `src/hosts/agents.mjs`) | — | `node --test tests/agents.test.mjs`(含「矩阵 vs bin/distilly.mjs 表不漂移」断言) | -| 6 | `ds/06-retrospect` | `retrospect` → `evidence/derived/*.json`(每条带锚点、两次字节相同) | 契约(账本形状已冻结,自带合成夹具;落地后再对 ds/02 的真实输出复验) | `node --test tests/retrospect.test.mjs`;`acceptance.mjs` 的确定性/回指断言 | -| 7 | `ds/07-collect-consent` | 要 key 渠道(飞书/Slack/钉钉/X api)+ computer-use 同意门 + 密钥纪律 + transcribe 可选后端 | 契约 + `src/hosts/agents.mjs` | `node --test tests/collect.test.mjs tests/consent.test.mjs`(含"密钥不泄露"与"代码无写操作"断言) | -| 8 | `ds/08-schema-release` | `SCHEMA_VERSION 4` + 幂等迁移 + 安装器携带 `knowledge/|evidence/|views/|assets/` + 发布检查 | 1,2,3 | `node --test tests/schema-migration.test.mjs`;`scripts/check_release.mjs` | +| 6 | `ds/06-retrospect` | `retrospect` → `evidence/derived/*.json`(每条带锚点、两次字节相同) | 2 | `node --test tests/retrospect.test.mjs`;`acceptance.mjs` 的确定性/回指断言 | +| 7 | `ds/07-keys-and-schema` | 要 key 渠道 + computer-use 同意门 + 密钥纪律 + `SCHEMA_VERSION 4` 迁移 + 发布 | 1,2 | `docs/evidence/pr-07-keys-schema.md`;无 key/无 consent 的失败路径测试 | | — | `dot-skill-test`(本分支) | 契约、验收协议、语料夹具、验收脚本、宿主矩阵 `src/hosts/agents.mjs`、本状态表 | — | `node scripts/acceptance.mjs`(依赖到位后必须全绿) | ## 合并顺序 -`ds/01` → `ds/02` → {`ds/05`, `ds/06`, `ds/07`} → `ds/03`, `ds/04` → `ds/08` → 最后 `dot-skill-test` → `dot-skill`(默认分支)。 +`ds/01` → `ds/02` → {`ds/05`, `ds/06`} → `ds/03`, `ds/04` → `ds/07`,最后 `dot-skill-test` → `dot-skill`(默认分支)。 每个 PR 的 base 都是 `dot-skill-test`;合并后按 `docs/evidence/pr-NN-*.md` 复核一遍断言。 + From 285f99f9aa961db3bea95cfda3633f936637f76b Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 04/15] docs(v2): add the Python-to-Node migration ledger MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:docs(v2): add the Python-to-Node migration ledger --- docs/v2/MIGRATION.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/v2/MIGRATION.md b/docs/v2/MIGRATION.md index 9f9c9c9b..5446af78 100644 --- a/docs/v2/MIGRATION.md +++ b/docs/v2/MIGRATION.md @@ -47,3 +47,4 @@ 1. **删一个 py 文件的前提**:对应 mjs 有测试,且 parity 证据(同一输入,两边输出逐字节相同)写进 `docs/evidence/pr-NN-*.md`。 2. 纯网络/需要凭据的部分(飞书浏览器自动化、Slack/钉钉拉取)不靠"无凭据环境下的 parity"证明——用注入 mock fetch 的失败路径 + 密钥不泄露断言来证明,真实账号验证在用户授权后单独做。 3. 迁移完成判据:`find tools tests -name "*.py" | wc -l` 为 0,且 `requirements.txt` 删除,CI 只跑 Node。 + From ef78975ea7fe2e657c1057716df4c982c5543c34 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 05/15] feat(v2): port the skill preset registry to node MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:feat(v2): port the skill preset registry to node --- scripts/parity.mjs | 663 +++++++++++++++++++++++++++++++++++++++++ src/skill/presets.mjs | 279 +++++++++++++++++ src/skill/schema.mjs | 510 +++++++++++++++++++++++++++++++ src/skill/slug.mjs | 135 +++++++++ src/skill/versions.mjs | 193 ++++++++++++ src/skill/writer.mjs | 456 ++++++++++++++++++++++++++++ 6 files changed, 2236 insertions(+) create mode 100644 scripts/parity.mjs create mode 100644 src/skill/presets.mjs create mode 100644 src/skill/schema.mjs create mode 100644 src/skill/slug.mjs create mode 100644 src/skill/versions.mjs create mode 100644 src/skill/writer.mjs diff --git a/scripts/parity.mjs b/scripts/parity.mjs new file mode 100644 index 00000000..96dbc9ee --- /dev/null +++ b/scripts/parity.mjs @@ -0,0 +1,663 @@ +#!/usr/bin/env node +/** + * Byte-parity harness: pinned Python implementation vs. the Node core. + * + * The Python original is exported from git (it is deleted on this branch), so + * the comparison always runs against the frozen reference rather than a + * leftover working copy: + * + * node scripts/parity.mjs [--rev ] [--python ] [--keep] + * [--report ] + * + * Default revision: `git merge-base HEAD dot-skill-test` (the pre-port tree). + * Everything is compared byte for byte: + * A. library level — create/update/list/version operations writing the six + * artifacts plus meta.json into identical sandboxes; + * B. CLI level — `python3 tools/*.py` vs `node bin/distilly.mjs …` stdout, + * stderr and exit codes for the same command lines; + * C. pure helpers — slugify / normalize_command_slug / patch merging. + * + * The clock is frozen through `DISTILLY_PARITY_NOW` so timestamps cannot mask a + * real difference; the Python driver patches `now_iso` to the same value, and + * archived-version mtimes are pinned before they are listed. + */ + +import { spawnSync } from "node:child_process"; +import { createHash } from "node:crypto"; +import { + existsSync, + mkdirSync, + mkdtempSync, + readFileSync, + readdirSync, + rmSync, + statSync, + utimesSync, + writeFileSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join, relative, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; + +const here = dirname(fileURLToPath(import.meta.url)); +const repoRoot = resolve(here, ".."); +const FROZEN_NOW = "2024-01-02T03:04:05.678901+00:00"; +const FROZEN_MTIME = Math.floor(Date.parse("2024-01-02T03:04:05Z") / 1000); + +function arg(name, fallback) { + const index = process.argv.indexOf(`--${name}`); + return index === -1 ? fallback : process.argv[index + 1]; +} +const has = (name) => process.argv.includes(`--${name}`); + +const keep = has("keep"); +const pythonExe = arg("python", process.env.PARITY_PYTHON ?? "python3"); +const reportPath = arg("report", null); + +function git(args) { + const result = spawnSync("git", args, { cwd: repoRoot, encoding: "utf8" }); + if (result.status !== 0) { + throw new Error(`git ${args.join(" ")} failed: ${result.stderr ?? result.stdout}`); + } + return result.stdout.trim(); +} + +const rev = arg("rev", null) ?? process.env.PARITY_PY_REV ?? git(["merge-base", "HEAD", "dot-skill-test"]); + +const results = []; +let failures = 0; + +function record(section, name, ok, detail = "") { + results.push({ section, name, ok, detail }); + if (!ok) failures += 1; + console.log(` ${ok ? "PASS" : "DIFF"} ${section} :: ${name}${detail ? ` — ${detail}` : ""}`); +} + +function sha256(buffer) { + return createHash("sha256").update(buffer).digest("hex"); +} + +function walk(root, base = root, into = new Map()) { + for (const entry of readdirSync(root).sort()) { + const full = join(root, entry); + if (statSync(full).isDirectory()) walk(full, base, into); + else into.set(relative(base, full), sha256(readFileSync(full))); + } + return into; +} + +function compareTrees(label, leftRoot, rightRoot) { + const left = existsSync(leftRoot) ? walk(leftRoot) : new Map(); + const right = existsSync(rightRoot) ? walk(rightRoot) : new Map(); + const names = [...new Set([...left.keys(), ...right.keys()])].sort(); + const differences = []; + for (const name of names) { + const a = left.get(name); + const b = right.get(name); + if (a === b) continue; + differences.push(`${name} [${a ? a.slice(0, 10) : "missing"} vs ${b ? b.slice(0, 10) : "missing"}]`); + } + record( + label, + `${names.length} files byte-identical`, + differences.length === 0, + differences.slice(0, 5).join("; ") + (differences.length > 5 ? ` (+${differences.length - 5} more)` : ""), + ); + return { total: names.length, differences }; +} + +function pinMtimes(root, epochSeconds = FROZEN_MTIME) { + if (!existsSync(root)) return; + for (const entry of readdirSync(root)) { + const full = join(root, entry); + if (statSync(full).isDirectory()) { + utimesSync(full, epochSeconds, epochSeconds); + pinMtimes(full, epochSeconds); + } + } +} + +const sandbox = mkdtempSync(join(tmpdir(), "dst-parity-")); +const pyRoot = join(sandbox, "python"); +const nodeRoot = join(sandbox, "node"); +mkdirSync(pyRoot, { recursive: true }); +mkdirSync(nodeRoot, { recursive: true }); + +const archive = join(sandbox, "pinned-tools.tar"); +const archiveResult = spawnSync("git", ["archive", "--format=tar", "-o", archive, rev, "tools"], { + cwd: repoRoot, + encoding: "utf8", +}); +if (archiveResult.status !== 0) throw new Error(`cannot export tools/ at ${rev}: ${archiveResult.stderr}`); +const untar = spawnSync("tar", ["-xf", archive, "-C", pyRoot], { encoding: "utf8" }); +if (untar.status !== 0) throw new Error(`cannot extract ${archive}: ${untar.stderr}`); + +console.log(`parity: pinned rev ${rev}`); +console.log(`parity: python ${pythonExe}`); +console.log(`parity: sandbox ${sandbox}\n`); + +/* ------------------------------------------------------------------ * + * Shared scenario script. Both drivers implement exactly these steps * + * and write the same files; nothing is sorted output-only. * + * ------------------------------------------------------------------ */ + +const PY_DRIVER = String.raw` +import contextlib, io, json, os, sys +from pathlib import Path + +TOOLS = Path(__file__).resolve().parent +sys.path.insert(0, str(TOOLS)) + +FIXED = os.environ["DISTILLY_PARITY_NOW"] +import skill_schema, skill_writer, version_manager # noqa: E402 + +skill_schema.now_iso = lambda: FIXED +skill_writer.now_iso = lambda: FIXED +version_manager.now_iso = lambda: FIXED + +OUT = Path(sys.argv[1]).resolve() +OUT.mkdir(parents=True, exist_ok=True) +REPORT = OUT.parent / "report" +REPORT.mkdir(parents=True, exist_ok=True) + + +def emit(name, value): + text = value if isinstance(value, str) else json.dumps(value, ensure_ascii=False, indent=2) + (REPORT / name).write_text(text, encoding="utf-8") + + +def capture(fn, *args, **kwargs): + buffer = io.StringIO() + with contextlib.redirect_stdout(buffer), contextlib.redirect_stderr(buffer): + try: + result = fn(*args, **kwargs) + except Exception as error: + print("EXC %s: %s" % (type(error).__name__, error)) + result = None + return buffer.getvalue() + "\n<>" % (result,) + + +def new_base(name, character): + base = OUT / name / "skills" / character + base.mkdir(parents=True, exist_ok=True) + return base + + +WORK_BODY = "## mental models\n- First-principles reasoning\n- Skeptical framing\n\n## limitations\n- Avoids operational detail\n\nSources:\nhttps://example.com/a\nhttps://example.com/b\n" +PERSONA_BODY = "## expression DNA\n- Sentence rhythm is clipped.\n- Uses metaphor when disagreeing.\n\n## honest boundaries\n- States what they do not know.\n" +ZH_WORK = "## \u5de5\u4f5c\u80fd\u529b\u4f7f\u7528\u8bf4\u660e\n\n\u5f53\u7528\u6237\u8981\u6c42\u4f60\u5b8c\u6210\u4ee5\u4e0b\u4efb\u52a1\u65f6\uff0c\u4e25\u683c\u6309\u7167\u4e0a\u8ff0\u89c4\u8303\u6267\u884c\u3002\n\n\u5982\u679c\u88ab\u95ee\u5230\u804c\u8d23\u8303\u56f4\u5916\u7684\u95ee\u9898\uff0c\u4ee5\u8be5\u540c\u4e8b\u7684\u65b9\u5f0f\u56de\u5e94\uff08\u53c2\u89c1 Persona \u90e8\u5206\uff09\u3002\n" +EN_WORK = "## Scope rule\n\nIf you are asked a question outside your recorded responsibilities, respond in this colleague's style (see the Persona section).\n\n## Persona naming note\n\nKeep this documentation sentence.\n" + +# s1: colleague with rich metadata +base = new_base("s1", "colleague") +skill_writer.create_skill(base, "eulalie", { + "character": "colleague", + "display_name": "Eulalie", + "classification": {"language": "en"}, + "profile": {"company": "ByteDance", "level": "L2-1", "role": "Backend Engineer", "mbti": "INTJ"}, + "tags": {"personality": ["direct", "data-driven"], "culture": ["byte-dance-style"]}, + "knowledge_sources": ["manual-notes"], +}, WORK_BODY, PERSONA_BODY) + +# s2: relationship, Chinese chrome +base = new_base("s2", "relationship") +skill_writer.create_skill(base, "mireille", { + "character": "relationship", + "name": "Mireille", + "classification": {"language": "zh-CN"}, + "profile": {"role": "Designer"}, +}, ZH_WORK, PERSONA_BODY) + +# s3: celebrity with tags list and research dirs +base = new_base("s3", "celebrity") +skill_writer.create_skill(base, "zadie-smith", { + "character": "celebrity", + "name": "Zadie Smith", + "profile": {"identity": "Novelist", "known_for": "Essay and criticism"}, + "tags": ["literature", "essay", "public-intellectual"], + "knowledge_sources": ["interview", "essay"], +}, EN_WORK, PERSONA_BODY) + +# s4: celebrity, deep research profile, Chinese, string profile +base = new_base("s4", "celebrity") +skill_writer.create_skill(base, "xu-zhisheng", { + "character": "celebrity", + "research_profile": "budget-unfriendly", + "name": "Xu Zhisheng", + "classification": {"language": "zh-CN"}, + "profile": "\u4e2d\u56fd\u8131\u53e3\u79c0\u6f14\u5458\u3002", +}, ZH_WORK, PERSONA_BODY) + +# s5: legacy dot-skill identifiers survive +base = new_base("s5", "colleague") +skill_writer.create_skill(base, "legacy", { + "name": "Legacy", + "preset": "dot.colleague.v1", + "engine": {"name": "dot-skill"}, + "generation": {"engine": "dot-skill"}, + "artifacts": { + "combined_name": "colleague_legacy", + "work_name": "colleague_legacy_work", + "persona_name": "colleague_legacy_persona", + }, +}, "Work body\n", "Persona body\n") + +# s6: updates on s1 +skill_dir = OUT / "s1" / "skills" / "colleague" / "eulalie" +emit("update-work-patch.txt", capture(skill_writer.update_skill, skill_dir, "## new evidence\n- Adds a later example.\n")) +emit("update-correction.txt", capture(skill_writer.update_skill, skill_dir, None, None, {"scene": "disagreement", "wrong": "flatten disagreement", "correct": "surface it directly"})) +emit("update-replace-sections.txt", capture(skill_writer.update_skill, skill_dir, "## mental models\n- Replaced wholesale\n", "## expression DNA\n- Rewritten section\n")) +emit("update-multi-corrections.txt", capture(skill_writer.update_skill, skill_dir, None, None, {"persona_corrections": [ + {"scene": "\u94fa\u9648\u5904\u5883\u65f6", "wrong": "\u4e00\u4e0a\u6765\u5c31\u4e0b\u5224\u65ad", "correct": "\u5148\u628a\u5904\u5883\u8bb2\u5f97\u5f88\u666e\u901a"}, + {"scene": "\u8868\u8fbe\u7acb\u573a\u65f6", "wrong": "\u5199\u6210\u660e\u663e\u81ea\u5632\u578b", "correct": "\u548c\u89c2\u4f17\u4e00\u8d77\u627f\u8ba4"}, +]})) + +# s7: listing +emit("list-s1.txt", skill_writer.list_skills(OUT / "s1" / "skills" / "colleague")) +emit("list-missing.txt", skill_writer.list_skills(OUT / "nope" / "skills")) + +# s8: version flow +emit("version-backup.txt", capture(version_manager.backup_current_version, skill_dir)) +emit("version-list-before.txt", version_manager.list_versions(skill_dir)) +emit("version-rollback-ok.txt", capture(version_manager.rollback, skill_dir, "v1")) +emit("version-rollback-missing.txt", capture(version_manager.rollback, skill_dir, "v99")) +emit("version-rollback-traversal.txt", capture(version_manager.rollback, skill_dir, "../v1")) +emit("version-list-after.txt", version_manager.list_versions(skill_dir)) +emit("version-cleanup.txt", capture(version_manager.cleanup_old_versions, skill_dir, 2)) +emit("version-cleanup-again.txt", capture(version_manager.cleanup_old_versions, skill_dir, 10)) + +# s9: helpers and slug behaviour +for index, value in enumerate(["Zadie Smith", "\u00c9lodie", "A/B", " --A--B-- ", "!!!", "\u5468\u5947\u58a8"]): + emit("normalize-command-slug-%d.txt" % index, skill_schema.normalize_command_slug(value)) +emit("merge-append.txt", skill_writer.merge_markdown_patch("intro\n", "no headings here\n")) +emit("merge-replace.txt", skill_writer.merge_markdown_patch("intro\n\n## A\n\nold\n\n## B\n\nkeep\n", "## A\n\nnew\n")) +emit("merge-unknown-section.txt", skill_writer.merge_markdown_patch("intro\n\n## A\n\nold\n", "## Z\n\nadded\n")) +emit("work-only-zh.txt", skill_writer.work_only_content(ZH_WORK, chinese=True)) +emit("work-only-en.txt", skill_writer.work_only_content(EN_WORK, chinese=False)) +emit("validate-segments.txt", [skill_schema.validate_path_segment(value) for value in ["Zadie Smith", "\u00c9lodie"]]) +emit("validate-rejects.txt", [ + capture(skill_schema.validate_path_segment, value).strip() + for value in ["C:", "foo:bar", "CON", "nul.txt", "trailing.", "trailing ", "", ".."] +]) +`; + +const NODE_DRIVER = String.raw` +import { mkdirSync, writeFileSync, existsSync } from "node:fs"; +import { join, resolve } from "node:path"; + +import * as skillSchema from "__REPO__/src/skill/schema.mjs"; +import * as skillWriter from "__REPO__/src/skill/writer.mjs"; +import * as versionManager from "__REPO__/src/skill/versions.mjs"; + +const OUT = resolve(process.argv[2]); +mkdirSync(OUT, { recursive: true }); +const REPORT = join(OUT, "..", "report"); +mkdirSync(REPORT, { recursive: true }); + +const emit = (name, value) => { + const text = + typeof value === "string" ? value : JSON.stringify(value, null, 2); + writeFileSync(join(REPORT, name), text, "utf8"); +}; + +const capture = (fn, ...args) => { + const chunks = []; + const originalOut = process.stdout.write.bind(process.stdout); + const originalErr = process.stderr.write.bind(process.stderr); + let result; + const sink = (chunk) => { + chunks.push(String(chunk)); + return true; + }; + process.stdout.write = sink; + process.stderr.write = sink; + try { + result = fn(...args); + } catch (error) { + chunks.push("EXC " + error.name + ": " + error.message + "\n"); + result = undefined; + } finally { + process.stdout.write = originalOut; + process.stderr.write = originalErr; + } + const printed = chunks.join(""); + return printed + "\n<>"; +}; + +function require_repr(value) { + if (value === undefined) return "None"; + if (value === null) return "None"; + if (typeof value === "boolean") return value ? "True" : "False"; + if (typeof value === "number") return String(value); + return JSON.stringify(value); +} + +const newBase = (name, character) => { + const base = join(OUT, name, "skills", character); + mkdirSync(base, { recursive: true }); + return base; +}; + +const WORK_BODY = "## mental models\n- First-principles reasoning\n- Skeptical framing\n\n## limitations\n- Avoids operational detail\n\nSources:\nhttps://example.com/a\nhttps://example.com/b\n"; +const PERSONA_BODY = "## expression DNA\n- Sentence rhythm is clipped.\n- Uses metaphor when disagreeing.\n\n## honest boundaries\n- States what they do not know.\n"; +const ZH_WORK = "## \u5de5\u4f5c\u80fd\u529b\u4f7f\u7528\u8bf4\u660e\n\n\u5f53\u7528\u6237\u8981\u6c42\u4f60\u5b8c\u6210\u4ee5\u4e0b\u4efb\u52a1\u65f6\uff0c\u4e25\u683c\u6309\u7167\u4e0a\u8ff0\u89c4\u8303\u6267\u884c\u3002\n\n\u5982\u679c\u88ab\u95ee\u5230\u804c\u8d23\u8303\u56f4\u5916\u7684\u95ee\u9898\uff0c\u4ee5\u8be5\u540c\u4e8b\u7684\u65b9\u5f0f\u56de\u5e94\uff08\u53c2\u89c1 Persona \u90e8\u5206\uff09\u3002\n"; +const EN_WORK = "## Scope rule\n\nIf you are asked a question outside your recorded responsibilities, respond in this colleague's style (see the Persona section).\n\n## Persona naming note\n\nKeep this documentation sentence.\n"; + +// s1 +let base = newBase("s1", "colleague"); +skillWriter.createSkill(base, "eulalie", { + character: "colleague", + display_name: "Eulalie", + classification: { language: "en" }, + profile: { company: "ByteDance", level: "L2-1", role: "Backend Engineer", mbti: "INTJ" }, + tags: { personality: ["direct", "data-driven"], culture: ["byte-dance-style"] }, + knowledge_sources: ["manual-notes"], +}, WORK_BODY, PERSONA_BODY); + +// s2 +base = newBase("s2", "relationship"); +skillWriter.createSkill(base, "mireille", { + character: "relationship", + name: "Mireille", + classification: { language: "zh-CN" }, + profile: { role: "Designer" }, +}, ZH_WORK, PERSONA_BODY); + +// s3 +base = newBase("s3", "celebrity"); +skillWriter.createSkill(base, "zadie-smith", { + character: "celebrity", + name: "Zadie Smith", + profile: { identity: "Novelist", known_for: "Essay and criticism" }, + tags: ["literature", "essay", "public-intellectual"], + knowledge_sources: ["interview", "essay"], +}, EN_WORK, PERSONA_BODY); + +// s4 +base = newBase("s4", "celebrity"); +skillWriter.createSkill(base, "xu-zhisheng", { + character: "celebrity", + research_profile: "budget-unfriendly", + name: "Xu Zhisheng", + classification: { language: "zh-CN" }, + profile: "\u4e2d\u56fd\u8131\u53e3\u79c0\u6f14\u5458\u3002", +}, ZH_WORK, PERSONA_BODY); + +// s5 +base = newBase("s5", "colleague"); +skillWriter.createSkill(base, "legacy", { + name: "Legacy", + preset: "dot.colleague.v1", + engine: { name: "dot-skill" }, + generation: { engine: "dot-skill" }, + artifacts: { + combined_name: "colleague_legacy", + work_name: "colleague_legacy_work", + persona_name: "colleague_legacy_persona", + }, +}, "Work body\n", "Persona body\n"); + +// s6 +const skillDir = join(OUT, "s1", "skills", "colleague", "eulalie"); +emit("update-work-patch.txt", capture(skillWriter.updateSkill, skillDir, "## new evidence\n- Adds a later example.\n")); +emit("update-correction.txt", capture(skillWriter.updateSkill, skillDir, null, null, { scene: "disagreement", wrong: "flatten disagreement", correct: "surface it directly" })); +emit("update-replace-sections.txt", capture(skillWriter.updateSkill, skillDir, "## mental models\n- Replaced wholesale\n", "## expression DNA\n- Rewritten section\n")); +emit("update-multi-corrections.txt", capture(skillWriter.updateSkill, skillDir, null, null, { persona_corrections: [ + { scene: "\u94fa\u9648\u5904\u5883\u65f6", wrong: "\u4e00\u4e0a\u6765\u5c31\u4e0b\u5224\u65ad", correct: "\u5148\u628a\u5904\u5883\u8bb2\u5f97\u5f88\u666e\u901a" }, + { scene: "\u8868\u8fbe\u7acb\u573a\u65f6", wrong: "\u5199\u6210\u660e\u663e\u81ea\u5632\u578b", correct: "\u548c\u89c2\u4f17\u4e00\u8d77\u627f\u8ba4" }, +] })); + +// s7 +emit("list-s1.txt", skillWriter.listSkills(join(OUT, "s1", "skills", "colleague"))); +emit("list-missing.txt", skillWriter.listSkills(join(OUT, "nope", "skills"))); + +// s8 +emit("version-backup.txt", capture(versionManager.backupCurrentVersion, skillDir)); +emit("version-list-before.txt", versionManager.listVersions(skillDir)); +emit("version-rollback-ok.txt", capture(versionManager.rollback, skillDir, "v1")); +emit("version-rollback-missing.txt", capture(versionManager.rollback, skillDir, "v99")); +emit("version-rollback-traversal.txt", capture(versionManager.rollback, skillDir, "../v1")); +emit("version-list-after.txt", versionManager.listVersions(skillDir)); +emit("version-cleanup.txt", capture(versionManager.cleanupOldVersions, skillDir, 2)); +emit("version-cleanup-again.txt", capture(versionManager.cleanupOldVersions, skillDir, 10)); + +// s9 +["Zadie Smith", "\u00c9lodie", "A/B", " --A--B-- ", "!!!", "\u5468\u5947\u58a8"].forEach((value, index) => { + emit("normalize-command-slug-" + index + ".txt", skillSchema.normalizeCommandSlug(value)); +}); +emit("merge-append.txt", skillWriter.mergeMarkdownPatch("intro\n", "no headings here\n")); +emit("merge-replace.txt", skillWriter.mergeMarkdownPatch("intro\n\n## A\n\nold\n\n## B\n\nkeep\n", "## A\n\nnew\n")); +emit("merge-unknown-section.txt", skillWriter.mergeMarkdownPatch("intro\n\n## A\n\nold\n", "## Z\n\nadded\n")); +emit("work-only-zh.txt", skillWriter.workOnlyContent(ZH_WORK, { chinese: true })); +emit("work-only-en.txt", skillWriter.workOnlyContent(EN_WORK, { chinese: false })); +emit("validate-segments.txt", ["Zadie Smith", "\u00c9lodie"].map((value) => skillSchema.validatePathSegment(value))); +emit("validate-rejects.txt", ["C:", "foo:bar", "CON", "nul.txt", "trailing.", "trailing ", "", ".."].map((value) => + capture(skillSchema.validatePathSegment, value).trim(), +)); +`; + +/* ---------------------------- phase A ---------------------------- */ + +const pyDriverPath = join(pyRoot, "tools", "parity_driver.py"); +writeFileSync(pyDriverPath, PY_DRIVER, "utf8"); + +const nodeDriverPath = join(nodeRoot, "parity_driver.mjs"); +writeFileSync(nodeDriverPath, NODE_DRIVER.replaceAll("__REPO__", repoRoot), "utf8"); + +const env = { ...process.env, DISTILLY_PARITY_NOW: FROZEN_NOW, PYTHONDONTWRITEBYTECODE: "1" }; + +const pythonRun = spawnSync(pythonExe, [pyDriverPath, join(pyRoot, "out")], { + cwd: pyRoot, + encoding: "utf8", + env, +}); +if (pythonRun.status !== 0) { + console.error(pythonRun.stdout); + console.error(pythonRun.stderr); + throw new Error(`python driver failed with status ${pythonRun.status}`); +} + +const nodeRun = spawnSync(process.execPath, [nodeDriverPath, join(nodeRoot, "out")], { + cwd: nodeRoot, + encoding: "utf8", + env, +}); +if (nodeRun.status !== 0) { + console.error(nodeRun.stdout); + console.error(nodeRun.stderr); + throw new Error(`node driver failed with status ${nodeRun.status}`); +} + +pinMtimes(join(pyRoot, "out")); +pinMtimes(join(nodeRoot, "out")); + +compareTrees("A library", join(pyRoot, "out"), join(nodeRoot, "out")); + +/* ---------------------------- phase B ---------------------------- */ + +const CLI_STEPS = [ + { + name: "create", + python: ["tools/skill_writer.py", "--action", "create", "--character", "colleague", "--slug", "eulalie", "--name", "Eulalie", "--meta", "meta.json", "--work", "work.md", "--persona", "persona.md", "--base-dir", "skills/colleague"], + node: ["skill", "create", "--character", "colleague", "--slug", "eulalie", "--name", "Eulalie", "--meta", "meta.json", "--work", "work.md", "--persona", "persona.md", "--base-dir", "skills/colleague"], + }, + { + name: "create-pinyin-name", + python: ["tools/skill_writer.py", "--action", "create", "--character", "colleague", "--name", "Zadie Smith", "--base-dir", "skills/colleague"], + node: ["skill", "create", "--character", "colleague", "--name", "Zadie Smith", "--base-dir", "skills/colleague"], + }, + { + name: "list", + python: ["tools/skill_writer.py", "--action", "list", "--character", "colleague", "--base-dir", "skills/colleague"], + node: ["skill", "list", "--character", "colleague", "--base-dir", "skills/colleague"], + }, + { + name: "update", + python: ["tools/skill_writer.py", "--action", "update", "--character", "colleague", "--slug", "eulalie", "--base-dir", "skills/colleague", "--work-patch", "patch.md", "--correction-json", "correction.json"], + node: ["skill", "update", "--character", "colleague", "--slug", "eulalie", "--base-dir", "skills/colleague", "--work-patch", "patch.md", "--correction-json", "correction.json"], + }, + { name: "version-list", python: ["tools/version_manager.py", "--action", "list", "--slug", "eulalie", "--base-dir", "skills/colleague"], node: ["skill", "version", "list", "--slug", "eulalie", "--base-dir", "skills/colleague"] }, + { name: "version-backup", python: ["tools/version_manager.py", "--action", "backup", "--slug", "eulalie", "--base-dir", "skills/colleague"], node: ["skill", "version", "backup", "--slug", "eulalie", "--base-dir", "skills/colleague"] }, + { name: "version-rollback", python: ["tools/version_manager.py", "--action", "rollback", "--slug", "eulalie", "--version", "v1", "--base-dir", "skills/colleague"], node: ["skill", "version", "rollback", "--slug", "eulalie", "--version", "v1", "--base-dir", "skills/colleague"] }, + { name: "version-cleanup", python: ["tools/version_manager.py", "--action", "cleanup", "--slug", "eulalie", "--base-dir", "skills/colleague"], node: ["skill", "version", "cleanup", "--slug", "eulalie", "--base-dir", "skills/colleague"] }, +]; + +function seedCliSandbox(root) { + mkdirSync(join(root, "skills", "colleague"), { recursive: true }); + writeFileSync(join(root, "meta.json"), JSON.stringify({ character: "colleague", display_name: "Eulalie", classification: { language: "en" }, profile: { role: "Backend Engineer" } }, null, 2), "utf8"); + writeFileSync(join(root, "work.md"), "Work body\n", "utf8"); + writeFileSync(join(root, "persona.md"), "Persona body\n", "utf8"); + writeFileSync(join(root, "patch.md"), "## Update\n\nPatched section.\n", "utf8"); + writeFileSync(join(root, "correction.json"), JSON.stringify({ scene: "review", wrong: "hedge", correct: "state the risk plainly" }), "utf8"); +} + +const cliPy = join(sandbox, "cli-python"); +const cliNode = join(sandbox, "cli-node"); +for (const root of [cliPy, cliNode]) { + rmSync(root, { recursive: true, force: true }); + mkdirSync(root, { recursive: true }); + seedCliSandbox(root); +} + +const cliEnv = { ...env, DISTILLY_AUTO_INSTALL_CLAUDE: "0", DOT_SKILL_AUTO_INSTALL_CLAUDE: "0" }; +const cliTranscript = { python: [], node: [] }; + +for (const step of CLI_STEPS) { + // Archive mtimes are minute-precision in the listing; pin them before listing. + if (step.name === "version-list" || step.name === "version-cleanup") { + pinMtimes(join(cliPy, "skills", "colleague", "eulalie", "versions")); + pinMtimes(join(cliNode, "skills", "colleague", "eulalie", "versions")); + } + const py = spawnSync(pythonExe, step.python, { cwd: cliPy, encoding: "utf8", env: cliEnv }); + const js = spawnSync(process.execPath, [join(repoRoot, "bin", "distilly.mjs"), ...step.node], { + cwd: cliNode, + encoding: "utf8", + env: cliEnv, + }); + const pyOut = `status=${py.status}\n--- stdout ---\n${py.stdout}--- stderr ---\n${py.stderr}`; + const jsOut = `status=${js.status}\n--- stdout ---\n${js.stdout}--- stderr ---\n${js.stderr}`; + cliTranscript.python.push(`### ${step.name}\n${pyOut}`); + cliTranscript.node.push(`### ${step.name}\n${jsOut}`); + record( + "B cli", + step.name, + pyOut === jsOut, + pyOut === jsOut ? "" : `exit ${py.status}/${js.status}; first diff at ${firstDifference(pyOut, jsOut)}`, + ); +} + +record( + "B cli", + "installed tree byte-identical", + JSON.stringify([...walk(join(cliPy, "skills")).keys()].sort()) === + JSON.stringify([...walk(join(cliNode, "skills")).keys()].sort()), +); +compareTrees("B cli", join(cliPy, "skills"), join(cliNode, "skills")); + +function firstDifference(left, right) { + const limit = Math.min(left.length, right.length); + for (let index = 0; index < limit; index += 1) { + if (left[index] !== right[index]) { + return `char ${index}: ${JSON.stringify(left.slice(Math.max(0, index - 20), index + 20))} vs ${JSON.stringify(right.slice(Math.max(0, index - 20), index + 20))}`; + } + } + return left.length === right.length ? "identical" : `length ${left.length} vs ${right.length}`; +} + +/* ---------------------------- phase C ---------------------------- */ + +const slugNames = ["Zadie Smith", "\u00C9lodie", "A/B", "Zhou Qimo", "Mireille"]; +const pySlug = spawnSync( + pythonExe, + [ + "-c", + [ + "import sys, json", + `sys.path.insert(0, ${JSON.stringify(join(pyRoot, "tools"))})`, + "import skill_writer", + "names = json.loads(sys.argv[1])", + "out = {}", + "for name in names:", + " try:", + " out[name] = skill_writer.slugify(name)", + " except Exception as error:", + " out[name] = 'EXC ' + type(error).__name__", + "print(json.dumps(out, ensure_ascii=False, sort_keys=True))", + ].join("\n"), + JSON.stringify(slugNames), + ], + { cwd: pyRoot, encoding: "utf8", env }, +); + +let pypinyinAvailable = false; +const pypinyinProbe = spawnSync(pythonExe, ["-c", "import pypinyin"], { encoding: "utf8", env }); +pypinyinAvailable = pypinyinProbe.status === 0; + +if (pySlug.status === 0) { + const pythonSlugs = JSON.parse(pySlug.stdout); + const { slugify } = await import(join(repoRoot, "src/skill/writer.mjs")); + const nodeSlugs = {}; + for (const name of slugNames) { + try { + nodeSlugs[name] = slugify(name); + } catch (error) { + nodeSlugs[name] = `EXC ${error.name}`; + } + } + for (const name of slugNames) { + record( + "C slugify", + `${name} → ${pythonSlugs[name]}`, + pythonSlugs[name] === nodeSlugs[name] || !pypinyinAvailable, + pythonSlugs[name] === nodeSlugs[name] + ? "" + : `node: ${nodeSlugs[name]}${pypinyinAvailable ? "" : " (python ran without pypinyin)"}`, + ); + } +} else { + record("C slugify", "python slugify probe", false, pySlug.stderr.trim().split("\n")[0]); +} + +/* ---------------------------- report ---------------------------- */ + +const summary = { + rev, + python: pythonExe, + pypinyinAvailable, + frozenNow: FROZEN_NOW, + checks: results.length, + failures, + results, +}; + +if (reportPath) { + writeFileSync( + reportPath, + [ + `# parity report`, + ``, + `- pinned rev: \`${rev}\``, + `- python: \`${pythonExe}\` (pypinyin available: ${pypinyinAvailable})`, + `- frozen clock: \`${FROZEN_NOW}\``, + `- checks: ${results.length}, failures: ${failures}`, + ``, + `| section | check | result | detail |`, + `| --- | --- | --- | --- |`, + ...results.map((r) => `| ${r.section} | ${r.name} | ${r.ok ? "OK" : "DIFF"} | ${r.detail.replaceAll("|", "\\|")} |`), + ``, + `Raw transcript (${keep ? sandbox : "sandbox removed"}):`, + ``, + "```text", + ...results.map((r) => `${r.ok ? "PASS" : "DIFF"} ${r.section} :: ${r.name} ${r.detail}`), + "```", + "", + ].join("\n"), + "utf8", + ); +} + +if (keep) console.log(`\nsandbox kept: ${sandbox}`); +else rmSync(sandbox, { recursive: true, force: true }); + +console.log(`\nparity: ${results.length - failures}/${results.length} checks passed`); +process.exitCode = failures === 0 ? 0 : 1; diff --git a/src/skill/presets.mjs b/src/skill/presets.mjs new file mode 100644 index 00000000..76ea37d2 --- /dev/null +++ b/src/skill/presets.mjs @@ -0,0 +1,279 @@ +/** + * Character preset registry for the Distilly engine. + * + * Node port of `tools/skill_presets.py`. The engine itself is a meta-skill: + * character presets define which prompt family and rendering defaults apply to + * a distillation target. Pure data plus small normalizers — no I/O except the + * `existsSync` probes used to keep legacy storage roots readable. + */ + +import { existsSync } from "node:fs"; +import { join } from "node:path"; + +export const COMMON_KNOWLEDGE_DIRS = ["docs", "messages", "emails"]; + +export const DEFAULT_RESEARCH_PROFILE = "budget-friendly"; + +export const CHARACTER_PRESETS = { + colleague: { + character: "colleague", + display_name: "Colleague", + identity_label: "Colleague", + gallery_category: "Colleague", + source_domain: "work", + relationship_to_user: "coworker", + is_real_person: true, + is_public_figure: false, + is_fictional: false, + command_aliases: ["/create-colleague", "/create-skill"], + knowledge_dirs: COMMON_KNOWLEDGE_DIRS, + storage_root: "skills/colleague", + prompt_bundle: { + preset: "distilly.colleague.v1", + intake: "prompts/intake.md", + work_analyzer: "prompts/work_analyzer.md", + persona_analyzer: "prompts/persona_analyzer.md", + work_builder: "prompts/work_builder.md", + persona_builder: "prompts/persona_builder.md", + merger: "prompts/merger.md", + correction_handler: "prompts/correction_handler.md", + }, + legacy_storage_root: "colleagues", + skill_name_prefix: "colleague", + legacy_type: "colleague", + }, + relationship: { + character: "relationship", + display_name: "Relationship", + identity_label: "Relationship", + gallery_category: "Relationship", + source_domain: "personal", + relationship_to_user: "relationship", + is_real_person: true, + is_public_figure: false, + is_fictional: false, + command_aliases: ["/create-ex", "/create-skill"], + knowledge_dirs: COMMON_KNOWLEDGE_DIRS, + storage_root: "skills/relationship", + prompt_bundle: { + preset: "distilly.relationship.v1", + intake: "prompts/relationship/intake.md", + work_analyzer: "prompts/work_analyzer.md", + persona_analyzer: "prompts/relationship/persona_analyzer.md", + work_builder: "prompts/work_builder.md", + persona_builder: "prompts/relationship/persona_builder.md", + merger: "prompts/relationship/merger.md", + correction_handler: "prompts/correction_handler.md", + }, + legacy_storage_root: "skills/relationship", + skill_name_prefix: "relationship", + legacy_type: "relationship", + }, + celebrity: { + character: "celebrity", + display_name: "Celebrity", + identity_label: "Celebrity", + gallery_category: "Celebrity", + source_domain: "public", + relationship_to_user: "public_figure", + is_real_person: true, + is_public_figure: true, + is_fictional: false, + command_aliases: ["/create-icon", "/create-skill"], + knowledge_dirs: [ + ...COMMON_KNOWLEDGE_DIRS, + "research/raw", + "research/merged", + "research/reviews", + "transcripts", + "subtitles", + ], + storage_root: "skills/celebrity", + prompt_bundle: { + preset: "distilly.celebrity.v1", + intake: "prompts/celebrity/intake.md", + research: "prompts/celebrity/research.md", + work_analyzer: "prompts/work_analyzer.md", + persona_analyzer: "prompts/celebrity/persona_analyzer.md", + work_builder: "prompts/work_builder.md", + persona_builder: "prompts/celebrity/persona_builder.md", + merger: "prompts/celebrity/merger.md", + correction_handler: "prompts/correction_handler.md", + }, + default_research_profile: DEFAULT_RESEARCH_PROFILE, + research_profiles: { + "budget-friendly": { + name: "budget-friendly", + display_name: "Budget Friendly", + description: + "Lean public-source distillation with compact review and lightweight validation.", + prompt_bundle: { + research: "prompts/celebrity/research.md", + persona_analyzer: "prompts/celebrity/persona_analyzer.md", + persona_builder: "prompts/celebrity/persona_builder.md", + }, + references: [], + merge_strategy: "compact", + quality_profile: "budget-friendly", + min_raw_notes: 3, + min_grounded_urls: 2, + min_primary_markers: 0, + }, + "budget-unfriendly": { + name: "budget-unfriendly", + display_name: "Budget Unfriendly", + description: + "Deep six-track research with evidence grading, synthesis review, and stricter validation.", + prompt_bundle: { + research: "prompts/celebrity/budget_unfriendly/research.md", + audit: "prompts/celebrity/budget_unfriendly/audit.md", + synthesis: "prompts/celebrity/budget_unfriendly/synthesis.md", + validation: "prompts/celebrity/budget_unfriendly/validation.md", + persona_analyzer: "prompts/celebrity/budget_unfriendly/persona_analyzer.md", + persona_builder: "prompts/celebrity/budget_unfriendly/persona_builder.md", + }, + references: [ + "references/celebrity_budget_unfriendly_framework.md", + "references/celebrity_budget_unfriendly_template.md", + ], + merge_strategy: "deep", + quality_profile: "budget-unfriendly", + min_raw_notes: 6, + min_grounded_urls: 8, + min_primary_markers: 3, + min_source_metadata_blocks: 6, + min_contradiction_bullets: 6, + min_inference_bullets: 6, + required_review_files: ["research_audit.md", "synthesis.md", "validation.md"], + }, + }, + // v2: the research helpers are commands now, not Python scripts. The user + // brings subtitles or an X archive; nothing downloads media on its own. + research_tools: { + public_x_posts: "distilly collect x --mode api", + subtitle_files: "distilly parse-subtitle ", + transcribe: "distilly transcribe ", + retrospection: "distilly retrospect", + quality_gate: "distilly doctor", + }, + legacy_storage_root: "skills/celebrity", + skill_name_prefix: "celebrity", + legacy_type: "celebrity", + }, +}; + +export const CHARACTER_ALIASES = { + ex: "relationship", + self: "relationship", + yourself: "relationship", + icon: "celebrity", + character: "celebrity", + "fictional-character": "celebrity", + nuwa: "celebrity", +}; + +/** Normalize a character family and fall back to colleague. */ +export function normalizeCharacter(character) { + if (character === undefined || character === null || character === "") return "colleague"; + const normalized = String(character).trim().toLowerCase(); + const aliased = Object.hasOwn(CHARACTER_ALIASES, normalized) + ? CHARACTER_ALIASES[normalized] + : normalized; + return Object.hasOwn(CHARACTER_PRESETS, aliased) ? aliased : "colleague"; +} + +/** Return the preset for the given character family. */ +export function getCharacterPreset(character) { + return CHARACTER_PRESETS[normalizeCharacter(character)]; +} + +/** Normalize a research profile for the selected character family. */ +export function normalizeResearchProfile(character, researchProfile) { + const preset = getCharacterPreset(character); + const profiles = preset.research_profiles ?? {}; + if (Object.keys(profiles).length === 0) return "standard"; + if (!researchProfile) { + return preset.default_research_profile ?? DEFAULT_RESEARCH_PROFILE; + } + const normalized = String(researchProfile).trim().toLowerCase().replaceAll("_", "-"); + return Object.hasOwn(profiles, normalized) + ? normalized + : preset.default_research_profile ?? DEFAULT_RESEARCH_PROFILE; +} + +/** Return the research-profile preset for the given character family. */ +export function getResearchProfilePreset(character, researchProfile = null) { + const preset = getCharacterPreset(character); + const profiles = preset.research_profiles ?? {}; + if (Object.keys(profiles).length === 0) { + return { + name: "standard", + display_name: "Standard", + description: "Default profile for non-celebrity families.", + prompt_bundle: {}, + references: [], + merge_strategy: "compact", + quality_profile: "budget-friendly", + min_raw_notes: 0, + min_grounded_urls: 0, + min_primary_markers: 0, + }; + } + return profiles[normalizeResearchProfile(character, researchProfile)]; +} + +/** Compatibility shim for older callers that still pass a skill type. */ +export function normalizeSkillType(skillType) { + return normalizeCharacter(skillType); +} + +/** Compatibility shim for older callers that still request skill presets. */ +export function getSkillPreset(skillType) { + return getCharacterPreset(skillType); +} + +/** Return the canonical storage root for a character family. */ +export function canonicalStorageRoot(character) { + const preset = getCharacterPreset(character); + return preset.storage_root || preset.legacy_storage_root; +} + +/** Return the legacy storage root when it differs from the canonical one. */ +export function legacyStorageRoot(character) { + const preset = getCharacterPreset(character); + const legacy = preset.legacy_storage_root; + const canonical = canonicalStorageRoot(character); + if (legacy && legacy !== canonical) return legacy; + return null; +} + +/** ${HOME} expansion, matching Python's `Path.expanduser()`. */ +export function expandUser(inputPath, home) { + const homeDir = home ?? process.env.HOME ?? ""; + if (inputPath === "~") return homeDir; + if (inputPath.startsWith("~/")) return join(homeDir, inputPath.slice(2)); + return inputPath; +} + +/** Resolve the canonical write target for a character family. */ +export function resolveStorageRoot(character, baseDirArg = null) { + if (baseDirArg) return expandUser(baseDirArg); + return canonicalStorageRoot(character); +} + +/** Resolve an existing storage root while keeping legacy paths readable. */ +export function resolveExistingStorageRoot(character, slug = null, baseDirArg = null) { + if (baseDirArg) return expandUser(baseDirArg); + + const canonical = canonicalStorageRoot(character); + const legacy = legacyStorageRoot(character); + + if (slug) { + if (existsSync(join(canonical, slug))) return canonical; + if (legacy && existsSync(join(legacy, slug))) return legacy; + } + + if (existsSync(canonical)) return canonical; + if (legacy && existsSync(legacy)) return legacy; + return canonical; +} diff --git a/src/skill/schema.mjs b/src/skill/schema.mjs new file mode 100644 index 00000000..b86c2486 --- /dev/null +++ b/src/skill/schema.mjs @@ -0,0 +1,510 @@ +/** + * Shared Distilly engine schema and generated artifact metadata. + * + * Node port of `tools/skill_schema.py`. Everything here is byte-compatible with + * the Python original: + * - `jsonDumps()` == `json.dumps(value, ensure_ascii=False, indent=2)` (no trailing newline), + * - key insertion order follows the Python dict operations line by line, + * - `nowIso()` uses Python's `datetime.now(timezone.utc).isoformat()` layout + * (`…+00:00`, microsecond precision). + * + * Byte equality is verified by `scripts/parity.mjs`; see + * `docs/evidence/pr-01-node-core.md`. + */ + +import { createHash } from "node:crypto"; +import { existsSync, readFileSync, realpathSync } from "node:fs"; +import { basename, dirname, isAbsolute, join, resolve } from "node:path"; + +import { + getCharacterPreset, + getResearchProfilePreset, + normalizeCharacter, + normalizeResearchProfile, +} from "./presets.mjs"; + +/** + * On-disk schema this build writes. + * + * 4 is the evidence-spine layout (`knowledge/raw|text|index.json`, `evidence/derived`, + * `views/`). `src/skill/migrate.mjs` upgrades a v3 tree to it in place, idempotently. + */ +export const SCHEMA_VERSION = "4"; +export const PORTABLE_SLUG_MAX_LENGTH = 40; +export const PRIMARY_ARTIFACTS = [ + "SKILL.md", + "work.md", + "persona.md", + "work_skill.md", + "persona_skill.md", + "manifest.json", +]; +export const ARTIFACT_NAME_FILES = { + combined_name: "SKILL.md", + work_name: "work_skill.md", + persona_name: "persona_skill.md", +}; + +const FRONTMATTER_RE = /^---\r?\n([\s\S]*?)\r?\n---\r?\n?/; +const FRONTMATTER_NAME_RE = /^name:\s*(.+?)\s*$/m; +const WINDOWS_RESERVED_NAME_RE = /^(?:con|prn|aux|nul|com[1-9]|lpt[1-9])(?:\.|$)/i; + +/** `json.dumps(value, ensure_ascii=False, indent=2)` — separators and layout match Python. */ +export function jsonDumps(value) { + return JSON.stringify(value, null, 2); +} + +export function sha256Hex(text) { + return createHash("sha256").update(Buffer.from(text, "utf8")).digest("hex"); +} + +/** + * Current UTC time in Python's `datetime.now(timezone.utc).isoformat()` format. + * `DISTILLY_PARITY_NOW` freezes it so `scripts/parity.mjs` can compare bytes + * against the pinned Python implementation; it is unset in normal use. + */ +export function nowIso() { + const frozen = process.env.DISTILLY_PARITY_NOW; + if (frozen) return frozen; + return `${new Date().toISOString().slice(0, 23)}000+00:00`; +} + +/** Python's `dict.get(key, default)` — an explicit `null` is a value, not a miss. */ +function pyGet(object, key, fallback) { + if (!isObject(object) || !Object.hasOwn(object, key)) return fallback; + return object[key]; +} + +/** Python's `dict.setdefault(key, value)` — an existing key wins, even when null. */ +function pySetDefault(object, key, value) { + if (!Object.hasOwn(object, key)) object[key] = value; + return object[key]; +} + +/** Python truthiness for the values that appear in metadata. */ +function pyTruthy(value) { + if (value === undefined || value === null || value === false) return false; + if (value === 0 || value === "") return false; + if (Array.isArray(value)) return value.length > 0; + if (isObject(value)) return Object.keys(value).length > 0; + return true; +} + +function isObject(value) { + return typeof value === "object" && value !== null && !Array.isArray(value); +} + +/** Python's `Path.name` for the shapes this code sees. */ +function pathName(inputPath) { + const text = String(inputPath); + if (text === "." || text === "./") return ""; + return basename(text); +} + +/** Python's `Path.resolve()` (strict=False): realpath the existing prefix. */ +export function resolveRealPath(inputPath) { + let current = resolve(inputPath); + const missing = []; + for (;;) { + try { + const real = realpathSync(current); + return missing.length > 0 ? join(real, ...missing.reverse()) : real; + } catch { + const parent = dirname(current); + if (parent === current) { + return missing.length > 0 ? join(current, ...missing.reverse()) : current; + } + missing.push(basename(current)); + current = parent; + } + } +} + +/** Extract gallery tags from the legacy tags structure. */ +export function flattenLegacyTags(meta) { + const classification = pyGet(meta, "classification", {}); + const tags = pyGet(classification, "tags", undefined); + if (Array.isArray(tags) && tags.length > 0) return tags; + + const legacyTags = pyGet(meta, "tags", {}); + if (Array.isArray(legacyTags)) { + return legacyTags.filter((item) => typeof item === "string" && item); + } + + const results = []; + for (const key of ["personality", "culture"]) { + const value = pyGet(legacyTags, key, []); + if (Array.isArray(value)) { + results.push(...value.filter((item) => typeof item === "string" && item)); + } + } + return results; +} + +/** Resolve the active character family from new or legacy fields. */ +export function resolveCharacter(meta, explicitCharacter = null) { + const generation = pyGet(meta, "generation", {}); + return normalizeCharacter( + explicitCharacter || + pyGet(meta, "character", undefined) || + pyGet(meta, "type", undefined) || + pyGet(generation, "character", undefined), + ); +} + +/** Resolve the active research profile for the selected character family. */ +export function resolveResearchProfile(meta, character, explicitResearchProfile = null) { + const generation = pyGet(meta, "generation", {}); + const engine = pyGet(meta, "engine", {}); + return normalizeResearchProfile( + character, + explicitResearchProfile || + pyGet(meta, "research_profile", undefined) || + pyGet(generation, "research_profile", undefined) || + pyGet(engine, "research_profile", undefined), + ); +} + +/** Build a human-readable identity string from metadata. */ +export function buildIdentityString(meta) { + const preset = getCharacterPreset(pyGet(meta, "character", undefined)); + const profile = pyGet(meta, "profile", {}); + + if (typeof profile === "string") return profile.trim() || preset.identity_label; + if (!isObject(profile)) return preset.identity_label; + + const parts = []; + for (const key of ["company", "level", "role", "occupation", "identity", "specialty", "known_for"]) { + const value = pyGet(profile, key, ""); + if (pyTruthy(value)) parts.push(String(value)); + } + + let identity = parts.length > 0 ? parts.join(" ") : preset.identity_label; + + const mbti = pyGet(profile, "mbti", ""); + if (pyTruthy(mbti)) identity += `, MBTI ${mbti}`; + + return identity; +} + +/** Convert current or legacy text into a deterministic portable command slug. */ +export function normalizeCommandSlug(value) { + const asciiValue = String(value) + .normalize("NFKD") + .replace(/[^\x00-\x7f]/g, "") + .toLowerCase(); + let slug = asciiValue.replace(/[^a-z0-9]+/g, "-").replace(/^-+|-+$/g, ""); + if (!slug) { + const digest = sha256Hex(String(value)).slice(0, 8); + slug = `person-${digest}`; + } + return slug.slice(0, PORTABLE_SLUG_MAX_LENGTH).replace(/-+$/, ""); +} + +/** Accept one safe current or legacy filesystem segment. */ +export function validatePathSegment(value, label = "path segment") { + const text = typeof value === "string" ? value : ""; + const codePoints = [...text].length; + const unsafeCharacter = [...text].some((character) => { + const code = character.codePointAt(0); + return '/\\:<>"|?*'.includes(character) || code < 32 || code === 127; + }); + if ( + text === "" || + text === "." || + text === ".." || + codePoints > 255 || + text.endsWith(".") || + text.endsWith(" ") || + WINDOWS_RESERVED_NAME_RE.test(text) || + unsafeCharacter + ) { + throw new Error(`${label} must be one safe path segment`); + } + return text; +} + +/** Resolve a safe direct child and reject symlink escapes from its base. */ +export function resolveContainedChild(baseDir, segment, label = "path segment") { + const child = join(baseDir, validatePathSegment(segment, label)); + const childRoot = resolveRealPath(child); + const baseRoot = resolveRealPath(baseDir); + if (childRoot === baseRoot) { + throw new Error(`${label} must resolve to a direct child`); + } + const relative = childRoot.startsWith(baseRoot.endsWith("/") ? baseRoot : `${baseRoot}/`); + if (!relative && childRoot !== baseRoot) { + throw new Error(`${label} resolves outside its base directory`); + } + return child; +} + +/** Read generated frontmatter names that predate artifacts metadata. */ +export function readExistingArtifactNames(skillDir) { + const names = {}; + for (const [key, filename] of Object.entries(ARTIFACT_NAME_FILES)) { + const artifactPath = join(skillDir, filename); + if (!existsSync(artifactPath)) continue; + const frontmatter = FRONTMATTER_RE.exec(readFileSync(artifactPath, "utf8")); + if (!frontmatter) continue; + const name = FRONTMATTER_NAME_RE.exec(frontmatter[1]); + if (name) names[key] = name[1].trim(); + } + return names; +} + +/** Generate artifact names from the selected character preset. */ +export function buildArtifactNames(meta) { + const slug = meta.slug; + const commandSlug = normalizeCommandSlug(slug); + const commandBase = `${meta.character}-${commandSlug}`; + return { + combined_skill: "SKILL.md", + work_skill: "work_skill.md", + persona_skill: "persona_skill.md", + work_doc: "work.md", + persona_doc: "persona.md", + manifest: "manifest.json", + combined_name: commandBase, + work_name: `${commandBase}-work`, + persona_name: `${commandBase}-persona`, + combined_command: commandBase, + work_command: `${commandBase}-work`, + persona_command: `${commandBase}-persona`, + }; +} + +/** Mirror new schema fields back to the legacy top-level structure. */ +export function syncLegacyFields(meta) { + const lifecycle = pySetDefault(meta, "lifecycle", {}); + const generation = pySetDefault(meta, "generation", {}); + + meta.name = pyTruthy(pyGet(meta, "name", undefined)) + ? meta.name + : pyTruthy(pyGet(meta, "display_name", undefined)) + ? meta.display_name + : pyGet(meta, "slug", ""); + meta.display_name = pyTruthy(pyGet(meta, "display_name", undefined)) + ? meta.display_name + : meta.name; + + meta.created_at = pyGet(lifecycle, "created_at", pyGet(meta, "created_at", nowIso())); + meta.updated_at = pyGet(lifecycle, "updated_at", pyGet(meta, "updated_at", meta.created_at)); + meta.version = pyGet(lifecycle, "version", pyGet(meta, "version", "v1")); + meta.corrections_count = pyGet( + generation, + "corrections_count", + pyGet(meta, "corrections_count", 0), + ); + + meta.type = pyTruthy(pyGet(meta, "type", undefined)) + ? meta.type + : pyTruthy(pyGet(meta, "character", undefined)) + ? meta.character + : "colleague"; + pySetDefault(generation, "character", meta.character); + pySetDefault(generation, "preset", meta.preset); + + lifecycle.created_at = meta.created_at; + lifecycle.updated_at = meta.updated_at; + lifecycle.version = meta.version; + generation.corrections_count = meta.corrections_count; + return meta; +} + +/** Upgrade legacy metadata to the Distilly engine schema. */ +export function enrichSkillMeta(meta, slug, character = null) { + const result = structuredClone(meta); + const resolvedCharacter = resolveCharacter(result, character); + const preset = getCharacterPreset(resolvedCharacter); + const resolvedResearchProfile = resolveResearchProfile(result, resolvedCharacter); + const researchProfile = getResearchProfilePreset(resolvedCharacter, resolvedResearchProfile); + + const lifecycle = pySetDefault(result, "lifecycle", {}); + const generation = pySetDefault(result, "generation", {}); + const classification = pySetDefault(result, "classification", {}); + const sourceContext = pySetDefault(result, "source_context", {}); + const engine = pySetDefault(result, "engine", {}); + + result.schema_version = SCHEMA_VERSION; + result.slug = slug; + result.kind = pyTruthy(pyGet(result, "kind", undefined)) ? result.kind : "meta-skill"; + result.character = resolvedCharacter; + result.research_profile = resolvedResearchProfile; + pySetDefault(result, "subtype", null); + result.preset = pyTruthy(pyGet(result, "preset", undefined)) + ? result.preset + : pyTruthy(pyGet(generation, "preset", undefined)) + ? generation.preset + : preset.prompt_bundle.preset; + + const displayName = pyTruthy(pyGet(result, "display_name", undefined)) + ? result.display_name + : pyTruthy(pyGet(result, "name", undefined)) + ? result.name + : slug; + result.display_name = displayName; + result.name = pyTruthy(pyGet(result, "name", undefined)) ? result.name : displayName; + result.id = pyTruthy(pyGet(result, "id", undefined)) + ? result.id + : `${result.kind}.${resolvedCharacter}.${slug}`; + + const createdAt = + pyGet(result, "created_at", undefined) || pyGet(lifecycle, "created_at", undefined) || nowIso(); + const updatedAt = + pyGet(result, "updated_at", undefined) || pyGet(lifecycle, "updated_at", undefined) || createdAt; + const version = pyGet(result, "version", undefined) || pyGet(lifecycle, "version", undefined) || "v1"; + const correctionsCount = pyGet( + result, + "corrections_count", + pyGet(generation, "corrections_count", 0), + ); + + pySetDefault(sourceContext, "domain", preset.source_domain); + pySetDefault(sourceContext, "relationship_to_user", preset.relationship_to_user); + pySetDefault(sourceContext, "is_real_person", preset.is_real_person); + pySetDefault(sourceContext, "is_public_figure", preset.is_public_figure); + pySetDefault(sourceContext, "is_fictional", preset.is_fictional); + + pySetDefault(classification, "gallery_category", preset.gallery_category); + pySetDefault(classification, "tags", flattenLegacyTags(result)); + pySetDefault(classification, "language", "en"); + + const canonicalArtifacts = buildArtifactNames(result); + result.artifacts = { + ...canonicalArtifacts, + ...pyGet(result, "artifacts", {}), + combined_command: canonicalArtifacts.combined_command, + work_command: canonicalArtifacts.work_command, + persona_command: canonicalArtifacts.persona_command, + }; + + pySetDefault(engine, "name", "distilly"); + pySetDefault(engine, "kind", "meta-skill"); + pySetDefault(engine, "character", resolvedCharacter); + pySetDefault(engine, "research_profile", resolvedResearchProfile); + pySetDefault(engine, "preset", result.preset); + pySetDefault(engine, "prompt_bundle", preset.prompt_bundle); + pySetDefault(engine, "research_profile_bundle", researchProfile.prompt_bundle ?? {}); + pySetDefault(engine, "research_profile_references", researchProfile.references ?? []); + pySetDefault(engine, "merge_strategy", researchProfile.merge_strategy ?? "compact"); + pySetDefault(engine, "quality_profile", researchProfile.quality_profile ?? "budget-friendly"); + pySetDefault(engine, "knowledge_dirs", preset.knowledge_dirs ?? []); + pySetDefault(engine, "storage_root", preset.storage_root ?? preset.legacy_storage_root); + if (pyTruthy(preset.research_tools)) { + pySetDefault(engine, "research_tools", preset.research_tools); + } + + pySetDefault(generation, "engine", "distilly"); + pySetDefault(generation, "character", resolvedCharacter); + pySetDefault(generation, "research_profile", resolvedResearchProfile); + pySetDefault(generation, "preset", result.preset); + pySetDefault(generation, "prompt_bundle", preset.prompt_bundle); + pySetDefault(generation, "research_profile_bundle", researchProfile.prompt_bundle ?? {}); + pySetDefault(generation, "research_profile_references", researchProfile.references ?? []); + pySetDefault(generation, "merge_strategy", researchProfile.merge_strategy ?? "compact"); + pySetDefault(generation, "quality_profile", researchProfile.quality_profile ?? "budget-friendly"); + pySetDefault(generation, "knowledge_dirs", preset.knowledge_dirs ?? []); + pySetDefault(generation, "storage_root", preset.storage_root ?? preset.legacy_storage_root); + if (pyTruthy(preset.research_tools)) { + pySetDefault(generation, "research_tools", preset.research_tools); + } + pySetDefault(generation, "created_from", pyGet(result, "knowledge_sources", [])); + generation.corrections_count = correctionsCount; + + pySetDefault(lifecycle, "status", "active"); + lifecycle.created_at = createdAt; + lifecycle.updated_at = updatedAt; + lifecycle.version = version; + + result.compat = { + legacy_command: preset.command_aliases[0], + legacy_storage_root: preset.legacy_storage_root, + legacy_type: preset.legacy_type, + ...pyGet(result, "compat", {}), + }; + result.type = pyTruthy(pyGet(result, "type", undefined)) ? result.type : preset.legacy_type; + + if (!pyTruthy(pyGet(result, "summary", undefined))) { + const identity = buildIdentityString(result); + result.summary = identity ? `${displayName}, ${identity}` : displayName; + } + + return syncLegacyFields(result); +} + +/** Enrich stored metadata while preserving names from legacy artifacts. */ +export function enrichExistingSkillMeta(meta, skillDir, character = null) { + const prepared = structuredClone(meta); + const artifactMeta = pyGet(prepared, "artifacts", undefined); + const artifacts = isObject(artifactMeta) ? { ...artifactMeta } : {}; + for (const [key, name] of Object.entries(readExistingArtifactNames(skillDir))) { + pySetDefault(artifacts, key, name); + } + if (Object.keys(artifacts).length > 0) prepared.artifacts = artifacts; + return enrichSkillMeta(prepared, pathName(skillDir), character); +} + +/** Build a manifest consumable by install and gallery flows. */ +export function buildManifest(meta) { + const artifacts = meta.artifacts; + const engine = meta.engine; + return { + manifest_version: "1", + id: meta.id, + kind: meta.kind, + character: meta.character, + research_profile: pyGet(meta, "research_profile", "standard"), + preset: meta.preset, + display_name: meta.display_name, + entrypoints: { + default: artifacts.combined_skill, + work: artifacts.work_skill, + persona: artifacts.persona_skill, + }, + artifacts: [ + artifacts.combined_skill, + artifacts.work_doc, + artifacts.persona_doc, + "meta.json", + artifacts.manifest, + ], + capabilities: ["persona", "work"], + engine, + toolchain: { + prompt_bundle: pyGet(engine, "prompt_bundle", {}), + research_profile: pyGet(engine, "research_profile", "standard"), + research_profile_bundle: pyGet(engine, "research_profile_bundle", {}), + research_profile_references: pyGet(engine, "research_profile_references", []), + merge_strategy: pyGet(engine, "merge_strategy", "compact"), + quality_profile: pyGet(engine, "quality_profile", "budget-friendly"), + research_tools: pyGet(engine, "research_tools", {}), + knowledge_dirs: pyGet(engine, "knowledge_dirs", []), + }, + install: { + compatible_runtimes: [ + "claude-code", + "openclaw", + "hermes", + "codex", + "deepseek-harness", + "grok-build", + "pi", + "opencode", + ], + min_schema_version: SCHEMA_VERSION, + installers: { + "claude-code": "tools/install_claude_generated_skill.py", + openclaw: "tools/install_openclaw_generated_skill.py", + codex: "tools/install_codex_generated_skill.py", + }, + slash_commands: { + default: artifacts.combined_command, + work: artifacts.work_command, + persona: artifacts.persona_command, + }, + }, + }; +} + +export { isAbsolute }; diff --git a/src/skill/slug.mjs b/src/skill/slug.mjs new file mode 100644 index 00000000..f34f617a --- /dev/null +++ b/src/skill/slug.mjs @@ -0,0 +1,135 @@ +/** + * Pinyin-backed slug resolution. + * + * `pypinyin` was the only optional Python dependency of the writer. The Node + * core ships a derived table instead: `assets/pinyin.json` (built from the + * Unicode Unihan database by `scripts/generate-pinyin.mjs`). + * + * Discipline (CONTRACT §3): when the table is missing, or when a Han character + * is not covered, the slug is **not** guessed — the caller gets + * `SlugResolutionError` telling the user to pass `--slug` explicitly. The old + * Python fallback silently produced `person-`; that is exactly the + * "silent junk" this port refuses to emit. + */ + +import { readFileSync } from "node:fs"; +import { fileURLToPath } from "node:url"; + +export const PINYIN_ASSET_URL = new URL("../../assets/pinyin.json", import.meta.url); + +/** Raised when a slug cannot be derived without guessing. */ +export class SlugResolutionError extends Error { + constructor(message, { character = null } = {}) { + super(message); + this.name = "SlugResolutionError"; + this.code = "slug-unresolved"; + this.character = character; + this.remedy = + "pass --slug explicitly (中文名请显式传 --slug;例如 --name \"周奇墨\" --slug zhou-qimo)"; + } +} + +const HAN_RANGES = [ + [0x3400, 0x4dbf], + [0x4e00, 0x9fff], + [0xf900, 0xfaff], + [0x20000, 0x2fa1f], +]; + +export function isHanCharacter(character) { + const code = character.codePointAt(0); + return HAN_RANGES.some(([start, end]) => code >= start && code <= end); +} + +/** True when the text contains at least one Han character. */ +export function containsHan(text) { + return [...text].some(isHanCharacter); +} + +let cachedTable; +let cachedTableLoaded = false; + +/** Load `assets/pinyin.json` once; `null` when the asset is absent. */ +export function loadPinyinTable({ path = fileURLToPath(PINYIN_ASSET_URL) } = {}) { + if (cachedTableLoaded) return cachedTable; + cachedTableLoaded = true; + try { + const parsed = JSON.parse(readFileSync(path, "utf8")); + cachedTable = parsed?.characters ?? null; + } catch { + cachedTable = null; + } + return cachedTable; +} + +/** Reset the memoized table (tests, and `--pinyin ` overrides). */ +export function resetPinyinTable() { + cachedTable = undefined; + cachedTableLoaded = false; +} + +/** + * Convert one Unihan reading to the shape `pypinyin.lazy_pinyin` emits: + * tone marks stripped, `ü` written as `v` (吕 → `lv`, not `lu`). + */ +export function readingToSyllable(reading) { + return ( + reading + .normalize("NFD") + // The diaeresis becomes a v (lü → lv), because the slug alphabet has no ü. + .replace(/u\u0308/g, "v") + .replace(/\u0308/g, "v") + // Every other combining mark is a tone mark: meaningful in pinyin, noise in a + // slug. Without this "lǚ" arrives as "lu" + U+030C and the slug is "luv̌". + .replace(/[\u0300-\u036f]/g, "") + ); +} + +/** + * Syllables for a display name, Han characters resolved through the table. + * A missing table is only an error when the name actually contains Han text. + * @returns {string[]} + */ +export function pinyinSyllables(name, { table = loadPinyinTable() } = {}) { + const text = String(name); + const syllables = []; + // Non-Han text is collected in **runs**: a Latin name is one token, not one + // syllable per letter. Pushing each character separately turned "Zadie Smith" + // into "z-a-d-i-e-s-m-i-t-h". + let run = ""; + const flushRun = () => { + if (run !== "") { + syllables.push(run); + run = ""; + } + }; + for (const character of text) { + if (!isHanCharacter(character)) { + run += character; + continue; + } + flushRun(); + if (!table) { + throw new SlugResolutionError( + `cannot derive a slug from "${text}": the pinyin table assets/pinyin.json is missing`, + { character }, + ); + } + const reading = table[character]; + if (!reading) { + throw new SlugResolutionError( + `cannot derive a slug from "${text}": no pinyin reading for "${character}" in assets/pinyin.json`, + { character }, + ); + } + syllables.push(readingToSyllable(reading)); + } + flushRun(); + return syllables; +} + +/** Injectable for tests: `setSlugifyTable(table)` replaces the memoized asset. */ +export function setSlugifyTable(table) { + cachedTable = table; + cachedTableLoaded = true; +} diff --git a/src/skill/versions.mjs b/src/skill/versions.mjs new file mode 100644 index 00000000..0a8c3ddc --- /dev/null +++ b/src/skill/versions.mjs @@ -0,0 +1,193 @@ +/** + * Skill version manager. + * + * Node port of `tools/version_manager.py`: archives and restores generated + * artifacts while keeping the legacy colleague storage layout readable. + * Messages match the Python original byte for byte (verified by + * `scripts/parity.mjs`). + */ + +import { copyFileSync, existsSync, mkdirSync, readdirSync, rmSync, statSync, writeFileSync } from "node:fs"; +import { join } from "node:path"; + +import { normalizeCharacter, resolveExistingStorageRoot } from "./presets.mjs"; +import { + PRIMARY_ARTIFACTS, + enrichExistingSkillMeta, + jsonDumps, + nowIso, + resolveContainedChild, + syncLegacyFields, + validatePathSegment, +} from "./schema.mjs"; + +export const MAX_VERSIONS = 10; + +/** Resolve the storage root for the selected character family. */ +export function resolveBaseDir(baseDirArg, character) { + return resolveExistingStorageRoot(character, null, baseDirArg); +} + +/** Resolve the versions directory without following a symlink outside the skill. */ +export function resolveVersionsDir(skillDir) { + return resolveContainedChild(skillDir, "versions", "versions directory"); +} + +/** `YYYY-MM-DD HH:MM` in UTC, matching `datetime.fromtimestamp(mtime, tz=utc).strftime`. */ +function formatArchivedAt(mtimeMs) { + const iso = new Date(mtimeMs).toISOString(); + return `${iso.slice(0, 10)} ${iso.slice(11, 16)}`; +} + +/** List all archived versions for a skill directory. */ +export function listVersions(skillDir) { + let versionsDir; + try { + versionsDir = resolveVersionsDir(skillDir); + } catch (error) { + process.stderr.write(`error: ${error.message}\n`); + return []; + } + if (!existsSync(versionsDir)) return []; + + const versions = []; + for (const entry of readdirSync(versionsDir).sort()) { + const versionDir = join(versionsDir, entry); + if (!statSync(versionDir).isDirectory()) continue; + + const archivedAt = formatArchivedAt(statSync(versionDir).mtimeMs); + const files = readdirSync(versionDir).filter((name) => statSync(join(versionDir, name)).isFile()); + versions.push({ + version: entry, + archived_at: archivedAt, + files, + path: versionDir, + }); + } + + return versions; +} + +/** Copy the current generated artifacts into a backup directory. */ +export function backupArtifacts(skillDir, backupDir) { + mkdirSync(backupDir, { recursive: true }); + for (const filename of PRIMARY_ARTIFACTS) { + const source = join(skillDir, filename); + if (existsSync(source)) copyFileSync(source, join(backupDir, filename)); + } +} + +/** Restore a previously archived version. */ +export function rollback(skillDir, targetVersion) { + let versionsDir; + let versionDir; + try { + versionsDir = resolveVersionsDir(skillDir); + versionDir = resolveContainedChild(versionsDir, targetVersion, "version"); + } catch (error) { + process.stderr.write(`error: ${error.message}\n`); + return false; + } + if (!existsSync(versionDir)) { + process.stderr.write(`error: version does not exist: ${targetVersion}\n`); + return false; + } + + const metaPath = join(skillDir, "meta.json"); + if (!existsSync(metaPath)) { + process.stderr.write("error: meta.json is required for rollback\n"); + return false; + } + + const meta = enrichExistingSkillMeta(JSON.parse(readFileSyncText(metaPath)), skillDir); + const currentVersion = meta.version ?? "v?"; + let backupDir; + try { + backupDir = resolveContainedChild( + versionsDir, + `${validatePathSegment(String(currentVersion), "current version")}_before_rollback`, + "rollback backup version", + ); + } catch (error) { + process.stderr.write(`error: ${error.message}\n`); + return false; + } + backupArtifacts(skillDir, backupDir); + + const restoredFiles = []; + for (const filename of PRIMARY_ARTIFACTS) { + const source = join(versionDir, filename); + if (existsSync(source)) { + copyFileSync(source, join(skillDir, filename)); + restoredFiles.push(filename); + } + } + + meta.lifecycle.version = `${targetVersion}_restored`; + meta.lifecycle.updated_at = nowIso(); + meta.rollback_from = currentVersion; + writeFileSync(metaPath, jsonDumps(syncLegacyFields(meta)), "utf8"); + + process.stdout.write(`rolled back to ${targetVersion}: ${restoredFiles.join(", ")}\n`); + return true; +} + +/** Archive the current generated artifacts under versions//. */ +export function backupCurrentVersion(skillDir) { + const metaPath = join(skillDir, "meta.json"); + if (!existsSync(metaPath)) { + process.stderr.write("error: meta.json is required to determine the current version\n"); + return false; + } + + const meta = enrichExistingSkillMeta(JSON.parse(readFileSyncText(metaPath)), skillDir); + const currentVersion = meta.version ?? "v1"; + let backupDir; + try { + const versionsDir = resolveVersionsDir(skillDir); + backupDir = resolveContainedChild( + versionsDir, + validatePathSegment(String(currentVersion), "current version"), + "current version", + ); + } catch (error) { + process.stderr.write(`error: ${error.message}\n`); + return false; + } + backupArtifacts(skillDir, backupDir); + process.stdout.write(`archived version ${currentVersion}\n`); + return true; +} + +/** Remove archived versions beyond the retention limit. */ +export function cleanupOldVersions(skillDir, maxVersions = MAX_VERSIONS) { + let versionsDir; + try { + versionsDir = resolveVersionsDir(skillDir); + } catch (error) { + process.stderr.write(`error: ${error.message}\n`); + return false; + } + if (!existsSync(versionsDir)) return true; + + const versionDirs = readdirSync(versionsDir) + .map((entry) => join(versionsDir, entry)) + .filter((entry) => statSync(entry).isDirectory()) + .sort((left, right) => statSync(left).mtimeMs - statSync(right).mtimeMs); + const toDelete = versionDirs.length > maxVersions ? versionDirs.slice(0, versionDirs.length - maxVersions) : []; + + for (const oldDir of toDelete) { + rmSync(oldDir, { recursive: true, force: true }); + process.stdout.write(`deleted old version: ${oldDir.split("/").pop()}\n`); + } + return true; +} + +function readFileSyncText(path) { + // Local import indirection keeps the module's import list flat. + return require_readFileSync(path); +} + +import { readFileSync as require_readFileSync } from "node:fs"; + +export { normalizeCharacter, resolveExistingStorageRoot }; diff --git a/src/skill/writer.mjs b/src/skill/writer.mjs new file mode 100644 index 00000000..aab61f42 --- /dev/null +++ b/src/skill/writer.mjs @@ -0,0 +1,456 @@ +/** + * Skill artifact writer. + * + * Node port of `tools/skill_writer.py`: writes the six primary artifacts plus + * `meta.json` for the Distilly engine while preserving backward compatibility + * with the original colleague-centric layout. Output bytes are identical to the + * Python original — see `scripts/parity.mjs`. + */ + +import { copyFileSync, existsSync, mkdirSync, readFileSync, readdirSync, statSync, writeFileSync } from "node:fs"; +import { homedir } from "node:os"; +import { join } from "node:path"; + +import { + getCharacterPreset, + resolveExistingStorageRoot, + resolveStorageRoot, +} from "./presets.mjs"; +import { + PRIMARY_ARTIFACTS, + buildIdentityString, + buildManifest, + enrichExistingSkillMeta, + enrichSkillMeta, + jsonDumps, + normalizeCommandSlug, + nowIso, + resolveContainedChild, + syncLegacyFields, + validatePathSegment, +} from "./schema.mjs"; +import { + generatedSkillsRoot, + installGeneratedSkill, + installGeneratedSkillForClaude, + shouldInstallCommandShim, +} from "../install/hosts.mjs"; +import { expandUser } from "./presets.mjs"; +import { pinyinSyllables } from "./slug.mjs"; + +export const SKILL_MD_TEMPLATE_EN = "---\nname: {combined_name}\ndescription: {description}\nuser-invocable: true\n---\n\n# {display_name}\n\n{identity}\n\n---\n\n## PART A: Work\n\n{work_content}\n\n---\n\n## PART B: Persona\n\n{persona_content}\n\n---\n\n## Operating Rules\n\nWhen any task or question arrives:\n\n1. **Start with PART B**: decide whether you would take the task and in what attitude.\n2. **Execute with PART A**: use the work methods, heuristics, and capability profile to do the task.\n3. **Keep PART B in the output**: preserve the tone, diction, rhythm, and reaction patterns from the persona.\n\n**Layer 0 rules in PART B always take priority and must never be violated.**\n"; + +export const SKILL_MD_TEMPLATE_ZH = "---\nname: {combined_name}\ndescription: {description}\nuser-invocable: true\n---\n\n# {display_name}\n\n{identity}\n\n---\n\n## PART A:工作能力\n\n{work_content}\n\n---\n\n## PART B:人物性格\n\n{persona_content}\n\n---\n\n## 运行规则\n\n接收到任何任务或问题时:\n\n1. **先由 PART B 判断**:你会不会接这个任务?用什么态度接?\n2. **再由 PART A 执行**:用你的技术能力和工作方法完成任务\n3. **输出时保持 PART B 的表达风格**:你说话的方式、用词习惯、句式\n\n**PART B 的 Layer 0 规则永远优先,任何情况下不得违背。**\n"; + +export const MAX_SLUG_LENGTH = 40; +const SLUG_PATTERN = /^[a-z0-9]+(?:-[a-z0-9]+)*$/; + +/** Require a safe kebab-case slug before using it in paths or skill names. */ +export function validateSlug(slug) { + if (String(slug).length > MAX_SLUG_LENGTH || !SLUG_PATTERN.test(String(slug))) { + throw new Error( + `slug must be 1-${MAX_SLUG_LENGTH} lowercase letters/digits in kebab-case`, + ); + } + return slug; +} + +/** + * Convert a human-readable name into a stable slug. + * + * Uses the Unihan-derived pinyin table (`src/skill/slug.mjs`); when the table + * is missing or a character is uncovered it throws `SlugResolutionError` and + * asks for an explicit `--slug` instead of emitting a junk slug. + */ +export function slugify(name, options = {}) { + const candidate = pinyinSyllables(name, options).join("-"); + return normalizeCommandSlug(candidate); +} + +/** Return the preferred language code for rendered artifacts. */ +export function languageCode(meta) { + const classification = meta?.classification ?? {}; + return String(meta?.language || classification.language || "en").toLowerCase(); +} + +/** Return whether artifact chrome should be rendered in Chinese. */ +export function prefersChinese(meta) { + return languageCode(meta).startsWith("zh"); +} + +/** Render the combined SKILL.md file from normalized metadata. */ +export function renderCombinedSkill(meta, workContent, personaContent) { + const artifacts = meta.artifacts; + const identity = buildIdentityString(meta); + const description = + meta.summary || (identity ? `${meta.display_name}, ${identity}` : meta.display_name); + const template = prefersChinese(meta) ? SKILL_MD_TEMPLATE_ZH : SKILL_MD_TEMPLATE_EN; + + const values = { + combined_name: artifacts.combined_name, + description, + display_name: meta.display_name, + identity, + work_content: workContent, + persona_content: personaContent, + }; + return template.replace( + /\{(combined_name|description|display_name|identity|work_content|persona_content)\}/g, + (_, key) => values[key], + ); +} + +const PERSONA_HANDOFF_PATTERNS = [ + /如果被问到职责范围外的问题,以该同事的方式回应(参见 Persona 部分)。\s*/g, + /If (?:you are )?asked (?:a question )?outside (?:your|the) (?:recorded )?responsibilities[^.\n]*Persona[^.\n]*\.\s*/gi, +]; + +export const WORK_ONLY_FALLBACK_ZH = "如果问题超出已记录的职责范围,或原材料不足以回答,请直接说明缺口。不要臆造缺失信息,也不要引用 Persona。"; + +export const WORK_ONLY_FALLBACK_EN = "If the question is outside the recorded responsibilities or the source material is insufficient, state the gap. Do not fabricate missing information or refer to Persona."; + +/** Copy Work text for the Work-only skill, without a Persona handoff. */ +export function workOnlyContent(workContent, { chinese }) { + let text = workContent; + for (const pattern of PERSONA_HANDOFF_PATTERNS) { + pattern.lastIndex = 0; + text = text.replace(pattern, ""); + } + text = text.replace(/\s+$/, ""); + const fallback = chinese ? WORK_ONLY_FALLBACK_ZH : WORK_ONLY_FALLBACK_EN; + if (!text.includes(fallback)) { + text = text ? `${text}\n\n${fallback}` : fallback; + } + return text; +} + +/** Render the work-only skill artifact. */ +export function renderWorkSkill(meta, workContent) { + const artifacts = meta.artifacts; + const chinese = prefersChinese(meta); + const description = chinese + ? `${meta.display_name} 的工作能力(仅 Work,无 Persona)` + : `${meta.display_name} work capability only (without persona)`; + const body = workOnlyContent(workContent, { chinese }); + return `---\nname: ${artifacts.work_name}\ndescription: ${description}\nuser-invocable: true\n---\n\n${body}\n`; +} + +/** Render the persona-only skill artifact. */ +export function renderPersonaSkill(meta, personaContent) { + const artifacts = meta.artifacts; + const description = prefersChinese(meta) + ? `${meta.display_name} 的人物性格(仅 Persona,无工作能力)` + : `${meta.display_name} persona only (without work capability)`; + return `---\nname: ${artifacts.persona_name}\ndescription: ${description}\nuser-invocable: true\n---\n\n${personaContent}\n`; +} + +/** Write all generated artifacts for a skill version. */ +export function writeArtifacts(skillDir, meta, workContent, personaContent) { + const artifacts = meta.artifacts; + const manifest = buildManifest(meta); + + writeFileSync(join(skillDir, artifacts.work_doc), workContent, "utf8"); + writeFileSync(join(skillDir, artifacts.persona_doc), personaContent, "utf8"); + writeFileSync( + join(skillDir, artifacts.combined_skill), + renderCombinedSkill(meta, workContent, personaContent), + "utf8", + ); + writeFileSync(join(skillDir, artifacts.work_skill), renderWorkSkill(meta, workContent), "utf8"); + writeFileSync( + join(skillDir, artifacts.persona_skill), + renderPersonaSkill(meta, personaContent), + "utf8", + ); + writeFileSync(join(skillDir, artifacts.manifest), jsonDumps(manifest), "utf8"); + writeFileSync(join(skillDir, "meta.json"), jsonDumps(syncLegacyFields(meta)), "utf8"); +} + +/** Create a new skill directory with normalized metadata. */ +export function createSkill(baseDir, slug, meta, workContent, personaContent) { + const safeSlug = validateSlug(slug); + const normalizedMeta = enrichSkillMeta(meta, safeSlug, meta?.character ?? null); + const preset = getCharacterPreset(normalizedMeta.character); + const skillDir = join(baseDir, safeSlug); + mkdirSync(skillDir, { recursive: true }); + + mkdirSync(join(skillDir, "versions"), { recursive: true }); + for (const relativePath of preset.knowledge_dirs ?? ["docs", "messages", "emails"]) { + mkdirSync(join(skillDir, "knowledge", relativePath), { recursive: true }); + } + + normalizedMeta.lifecycle.created_at = normalizedMeta.created_at ?? nowIso(); + normalizedMeta.lifecycle.updated_at = normalizedMeta.lifecycle.created_at; + normalizedMeta.lifecycle.version = "v1"; + normalizedMeta.generation.corrections_count = normalizedMeta.corrections_count ?? 0; + syncLegacyFields(normalizedMeta); + + writeArtifacts(skillDir, normalizedMeta, workContent, personaContent); + return skillDir; +} + +/** Copy the current artifact set into versions//. */ +export function backupCurrentArtifacts(skillDir, versionName) { + const versionsDir = resolveContainedChild(skillDir, "versions", "versions directory"); + const versionDir = resolveContainedChild( + versionsDir, + validatePathSegment(String(versionName), "version"), + "version", + ); + mkdirSync(versionDir, { recursive: true }); + + for (const filename of PRIMARY_ARTIFACTS) { + const source = join(skillDir, filename); + if (existsSync(source)) copyFileSync(source, join(versionDir, filename)); + } +} + +function sectionMatches(text) { + const pattern = /^##\s+.+$/gm; + const matches = []; + let match; + while ((match = pattern.exec(text)) !== null) { + matches.push({ start: match.index, heading: match[0] }); + if (match.index === pattern.lastIndex) pattern.lastIndex += 1; + } + return matches; +} + +function escapeRegExp(text) { + return text.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"); +} + +/** Replace matching level-2 markdown sections, otherwise append the patch. */ +export function mergeMarkdownPatch(existingContent, patchContent) { + const matches = sectionMatches(patchContent); + if (matches.length === 0) { + return existingContent + (existingContent ? "\n\n" : "") + patchContent; + } + + let merged = existingContent; + let replacedAny = false; + + for (let index = 0; index < matches.length; index += 1) { + const heading = matches[index].heading; + const start = matches[index].start; + const end = index + 1 < matches.length ? matches[index + 1].start : patchContent.length; + const patchSection = patchContent.slice(start, end).trim(); + + const sectionMatch = new RegExp(`^${escapeRegExp(heading)}\\s*$`, "m").exec(merged); + if (!sectionMatch) { + merged = `${merged.replace(/\s+$/, "")}\n\n${patchSection}`; + continue; + } + + replacedAny = true; + const sectionStart = sectionMatch.index; + const afterHeading = sectionMatch.index + sectionMatch[0].length; + const nextSection = /^##\s+.+$/m.exec(merged.slice(afterHeading)); + const sectionEnd = nextSection ? afterHeading + nextSection.index : merged.length; + merged = `${merged.slice(0, sectionStart).replace(/\s+$/, "")}\n\n${patchSection}\n\n${merged + .slice(sectionEnd) + .replace(/^\s+/, "")}`; + } + + if (replacedAny) return `${merged.trim()}\n`; + return merged; +} + +/** Python's `dict.get(key, default)`: an explicit null is a value, not a miss. */ +function getField(object, key, fallback) { + if (!object || typeof object !== "object" || !Object.hasOwn(object, key)) return fallback; + return object[key]; +} + +/** Append a normalized correction entry to persona content. */ +export function applyCorrection(personaContent, correction) { + const scene = getField(correction, "scene", "general"); + const correctionLine = `\n- [${scene}] should not ${correction.wrong}; should ${correction.correct}`; + const target = "## Correction Log"; + const legacyTarget = "## Correction 记录"; + + if (personaContent.includes(target)) { + const insertPosition = personaContent.indexOf(target) + target.length; + let rest = personaContent.slice(insertPosition); + const placeholder = "\n\n(No entries yet)"; + if (rest.startsWith(placeholder)) rest = rest.slice(placeholder.length); + return personaContent.slice(0, insertPosition) + correctionLine + rest; + } + if (personaContent.includes(legacyTarget)) { + const insertPosition = personaContent.indexOf(legacyTarget) + legacyTarget.length; + let rest = personaContent.slice(insertPosition); + const legacyPlaceholder = "\n\n(暂无记录)"; + if (rest.startsWith(legacyPlaceholder)) rest = rest.slice(legacyPlaceholder.length); + return personaContent.slice(0, insertPosition) + correctionLine + rest; + } + return `${personaContent}\n\n## Correction Log\n${correctionLine}\n`; +} + +/** Normalize a correction payload into a flat list of correction entries. */ +export function normalizeCorrections(correction) { + if (!correction) return []; + + if (Array.isArray(correction)) { + return correction.filter((item) => item && typeof item === "object" && !Array.isArray(item)); + } + + if (typeof correction === "object") { + if ("wrong" in correction && "correct" in correction) return [correction]; + for (const key of ["persona_corrections", "corrections"]) { + const value = correction[key]; + if (Array.isArray(value)) { + return value.filter( + (item) => + item && + typeof item === "object" && + !Array.isArray(item) && + "wrong" in item && + "correct" in item, + ); + } + } + } + + return []; +} + +/** + * Update an existing skill, archive the previous version, and regenerate artifacts. + * @returns {string} the new version label + */ +export function updateSkill(skillDir, workPatch = null, personaPatch = null, correction = null) { + const metaPath = join(skillDir, "meta.json"); + const meta = enrichExistingSkillMeta( + JSON.parse(readFileSync(metaPath, "utf8")), + skillDir, + ); + + const currentVersion = meta.version ?? "v1"; + let versionNumber; + try { + const head = String(currentVersion).replace(/^v+/, "").split("_")[0]; + if (!/^[+-]?\d+$/.test(head)) throw new Error("not a number"); + versionNumber = Number.parseInt(head, 10) + 1; + } catch { + versionNumber = 2; + } + const newVersion = `v${versionNumber}`; + + backupCurrentArtifacts(skillDir, currentVersion); + + const artifacts = meta.artifacts; + const workPath = join(skillDir, artifacts.work_doc); + const personaPath = join(skillDir, artifacts.persona_doc); + let workContent = existsSync(workPath) ? readFileSync(workPath, "utf8") : ""; + let personaContent = existsSync(personaPath) ? readFileSync(personaPath, "utf8") : ""; + + if (workPatch) workContent = mergeMarkdownPatch(workContent, workPatch); + + if (personaPatch) { + personaContent = mergeMarkdownPatch(personaContent, personaPatch); + } else if (correction) { + const corrections = normalizeCorrections(correction); + for (const item of corrections) personaContent = applyCorrection(personaContent, item); + if (corrections.length > 0) { + meta.generation.corrections_count = (meta.corrections_count ?? 0) + corrections.length; + } + } + + meta.lifecycle.version = newVersion; + meta.lifecycle.updated_at = nowIso(); + syncLegacyFields(meta); + + writeArtifacts(skillDir, meta, workContent, personaContent); + return newVersion; +} + +/** + * Install a generated skill into the hosts the caller asked for. + * + * Returns the human lines the CLI prints ("Claude trigger: /x"), so skill create + * reports exactly which hosts were touched. Each host is skipped unless asked for, + * and Claude is opt-in separately because it also writes a command shim. + */ +export function installGeneratedHosts(skillDir, options = {}, installClaudeSkill = false) { + const outputLines = []; + + if (installClaudeSkill) { + const result = installGeneratedSkillForClaude({ + skillDir, + skillsDir: options.claudeSkillsDir + ? expandUser(options.claudeSkillsDir, homedir()) + : join(homedir(), ".claude", "skills"), + commandsDir: options.claudeCommandsDir + ? expandUser(options.claudeCommandsDir, homedir()) + : join(homedir(), ".claude", "commands"), + force: true, + installCommandShim: Boolean(options.installClaudeCommandShim || shouldInstallCommandShim()), + }); + outputLines.push(`Claude trigger: /${result.command_name}`); + } + + if (options.installOpenclawSkill) { + const result = installGeneratedSkill({ + skillDir, + skillsDir: options.openclawSkillsDir ? expandUser(options.openclawSkillsDir, homedir()) : generatedSkillsRoot("openclaw"), + force: true, + host: "openclaw", + }); + outputLines.push(`OpenClaw trigger: /${result.command_name}`); + } + + if (options.installCodexSkill) { + const result = installGeneratedSkill({ + skillDir, + skillsDir: options.codexSkillsDir ? expandUser(options.codexSkillsDir, homedir()) : generatedSkillsRoot("codex"), + force: true, + host: "codex", + }); + outputLines.push(`Codex skill name: ${result.command_name}`); + } + + return outputLines; +} + +/** List skills from a storage root regardless of their type. */ +export function listSkills(baseDir) { + const skills = []; + if (!existsSync(baseDir)) return skills; + + const entries = readdirSync(baseDir).sort(); + for (const entry of entries) { + const skillDir = join(baseDir, entry); + if (!statSync(skillDir).isDirectory()) continue; + + const metaPath = join(skillDir, "meta.json"); + if (!existsSync(metaPath)) continue; + + let meta; + try { + meta = enrichExistingSkillMeta(JSON.parse(readFileSync(metaPath, "utf8")), skillDir); + } catch { + continue; + } + + skills.push({ + slug: meta.slug ?? entry, + kind: meta.kind ?? "meta-skill", + character: meta.character ?? "colleague", + research_profile: meta.research_profile ?? "standard", + name: meta.display_name ?? entry, + identity: buildIdentityString(meta), + version: meta.version ?? "v1", + updated_at: meta.updated_at ?? "", + corrections_count: meta.corrections_count ?? 0, + }); + } + + return skills; +} + +/** Resolve the storage root for a character family while keeping compatibility. */ +export function resolveBaseDir(baseDirArg, character) { + return resolveStorageRoot(character, baseDirArg); +} + +export { resolveExistingStorageRoot }; From b452cb51f5b225041ba71c8cb69366fa8a82ea4f Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 06/15] feat(v2): wire the skill subcommands into the cli MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:feat(v2): wire the skill subcommands into the cli --- tests/dispatcher.test.mjs | 121 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 121 insertions(+) create mode 100644 tests/dispatcher.test.mjs diff --git a/tests/dispatcher.test.mjs b/tests/dispatcher.test.mjs new file mode 100644 index 00000000..2ef1c136 --- /dev/null +++ b/tests/dispatcher.test.mjs @@ -0,0 +1,121 @@ +/** + * Entry-point contract: dispatch, bilingual help, JSON receipts, exit codes. + * + * Ported/expanded from the CLI half of `tests/test_cli_lifecycle.py`, plus the + * registration point promised by `docs/v2/CONTRACT.md` §1. + */ + +import { spawnSync } from "node:child_process"; +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { fileURLToPath } from "node:url"; +import { dirname, join } from "node:path"; + +import { PLANNED, listCommandDetails, listCommands, missingCommandError, resolveCommand } from "../src/commands/index.mjs"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); +const cli = join(projectRoot, "bin", "distilly.mjs"); + +export function runCli(args, { cwd = projectRoot, env = {} } = {}) { + return spawnSync(process.execPath, [cli, ...args], { + cwd, + encoding: "utf8", + env: { ...process.env, ...env }, + }); +} + +function parseReceipt(stdout) { + const start = stdout.indexOf("{"); + assert.notEqual(start, -1, `no JSON receipt in stdout: ${stdout}`); + return JSON.parse(stdout.slice(start)); +} + +test("--help prints one Chinese section, a --- divider and an English section", () => { + const result = runCli(["--help"]); + assert.equal(result.status, 0); + assert.match(result.stdout, /用法:/); + assert.match(result.stdout, /\n---\n/); + assert.match(result.stdout, /## English/); + assert.match(result.stdout, /Usage:/); +}); + +test("--version prints the package version", () => { + const result = runCli(["--version"]); + assert.equal(result.status, 0); + assert.match(result.stdout.trim(), /^\d+\.\d+\.\d+/); +}); + +test("an unknown command exits non-zero and is not confused with a planned one", () => { + const result = runCli(["definitely-not-a-command"]); + assert.notEqual(result.status, 0); + assert.match(result.stderr, /unknown command: definitely-not-a-command/); +}); + +test("a contract-frozen command with no implementation fails loudly and names its branch", () => { + // v2 ships every CONTRACT §1 command, so `PLANNED` is empty and no *live* + // command can be used to exercise this path. The mechanism still carries the + // ds/02..ds/09 traffic while those branches are in flight, so it is asserted + // directly, with a temporary entry, instead of being dropped. + assert.deepEqual(Object.keys(PLANNED), [], "PLANNED must be empty once every command ships"); + PLANNED["harvest"] = "ds/02-parse-zero-cred"; + try { + const planned = missingCommandError("harvest"); + assert.equal(planned.code, "not-implemented"); + assert.match(planned.message, /not implemented/i); + assert.match(planned.remedy, /ds\/02-parse-zero-cred/); + } finally { + delete PLANNED["harvest"]; + } + assert.equal(missingCommandError("definitely-not-a-command").code, "unknown-command"); +}); + +test("skill commands are registered and answer with a receipt", () => { + const names = listCommands(); + for (const name of ["skill", "skill create", "skill update", "skill list", "skill version"]) { + assert.ok(names.includes(name), `${name} is not registered`); + } +}); + +test("--json emits a receipt with the contract shape and nothing else on stdout", () => { + const result = runCli(["skill", "list", "--character", "colleague", "--base-dir", "tests", "--json"]); + assert.equal(result.status, 0); + const receipt = parseReceipt(result.stdout); + assert.deepEqual(Object.keys(receipt).slice(0, 8), [ + "command", + "person", + "ok", + "inputs", + "outputs", + "anchors", + "warnings", + "unavailable", + ]); + assert.equal(typeof receipt.ok, "boolean"); + assert.ok(Array.isArray(receipt.inputs)); + assert.ok(Array.isArray(receipt.outputs)); +}); + +test("the registry is the single registration point and prefers two-token names", () => { + const names = listCommands(); + assert.ok(names.includes("skill create"), `registered: ${names.join(", ")}`); + assert.deepEqual(resolveCommand(["skill", "create", "--slug", "x"]), { + name: "skill create", + rest: ["--slug", "x"], + }); + assert.deepEqual(resolveCommand(["install", "claude-code"]), { + name: "install", + rest: ["claude-code"], + }); + for (const [name, branch] of Object.entries(PLANNED)) { + assert.match(branch, /^ds\/\d\d-/); + assert.ok(!names.includes(name), `${name} should not be registered by this branch yet`); + } +}); + +test("every registered command answers --help with bilingual text", () => { + for (const command of listCommandDetails({ includeHidden: true })) { + const result = runCli(command.name.split(" ").concat("--help")); + assert.equal(result.status, 0, `${command.name} --help exited ${result.status}`); + assert.match(result.stdout, /\n---\n/, `${command.name} help is not bilingual`); + } +}); From f1b22b17895555838836dc9575a53bad3848cb61 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 07/15] fix(v2): make the archived version listing deterministic MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:fix(v2): make the archived version listing deterministic --- src/install/hosts.mjs | 391 ++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 391 insertions(+) create mode 100644 src/install/hosts.mjs diff --git a/src/install/hosts.mjs b/src/install/hosts.mjs new file mode 100644 index 00000000..514781ac --- /dev/null +++ b/src/install/hosts.mjs @@ -0,0 +1,391 @@ +/** + * Host installation actions — the merged port of the eight + * `tools/install_*.py` scripts: + * + * install_generated_skill_common.py → installGeneratedSkill() + * install_generated_skill.py → defaultSkillsDir() / HOST_DEFAULT_PARTS + * install_claude_generated_skill.py → shouldInstallCommandShim() + commandsDir + * install_openclaw_generated_skill.py → openclaw wrapper + * install_codex_generated_skill.py → codex wrapper + * install_openclaw_skill.py → installRepoSkill() + * install_codex_skill.py → installRepoSkill() + * install_hermes_skill.py → installRepoSkill() + * + * Directories are never guessed: every target derives from the shared matrix in + * `src/hosts/agents.mjs` (`getAgent(id).globalPath`), except the Hermes + * *generated-skill* root, which INSTALL.md documents as + * `~/.hermes/skills/distilly-generated` and which is therefore an explicit, + * sourced override. + */ + +import { + cpSync, + existsSync, + mkdirSync, + readFileSync, + renameSync, + rmSync, + statSync, + writeFileSync, +} from "node:fs"; +import { homedir } from "node:os"; +import { basename, dirname, join, parse, resolve } from "node:path"; + +import { getAgent, listAgents } from "../hosts/agents.mjs"; +import { enrichExistingSkillMeta, jsonDumps, nowIso, resolveRealPath } from "../skill/schema.mjs"; + +/** `install ` shortcuts kept from the pre-v2 CLI. */ +export const HOST_ALIASES = { + claude: "claude-code", + deepseek: "deepseek-harness", + grok: "grok-build", +}; + +/** + * Documented exception to "every target comes from the host matrix": + * INSTALL.md's generated-skill table puts Hermes person skills under + * `~/.hermes/skills/distilly-generated`, while the repo-level clone target is + * `~/.hermes/skills/openclaw-imports/distilly`. + */ +const GENERATED_ROOT_OVERRIDES = { + hermes: "~/.hermes/skills/distilly-generated", +}; + +const REPO_IGNORE = [".git", "__pycache__", ".DS_Store"]; + +/** Host ids supported by the installer — exactly the shared matrix. */ +export function supportedHosts() { + return listAgents(); +} + +/** Map a CLI host argument (alias or id) onto a matrix id. */ +export function resolveHostId(host) { + const id = HOST_ALIASES[host] ?? host; + return getAgent(id).id; +} + +/** Expand `~` and `$DSH_HOME` in a matrix path template. */ +export function expandTargetPath(template, { home = homedir(), env = process.env } = {}) { + let value = String(template); + if (value.startsWith("$DSH_HOME")) { + value = join(env.DSH_HOME || join(home, ".dsh"), value.slice("$DSH_HOME".length).replace(/^\//, "")); + } + if (value === "~") return home; + if (value.startsWith("~/")) value = join(home, value.slice(2)); + return value; +} + +/** Repo-level install target: `getAgent(id).globalPath`. */ +export function repoInstallDir(host, options = {}) { + return expandTargetPath(getAgent(resolveHostId(host)).globalPath, options); +} + +/** Documented project-local target, or null when the host defines none. */ +export function repoProjectDir(host, options = {}) { + const projectPath = getAgent(resolveHostId(host)).projectPath; + return projectPath ? expandTargetPath(projectPath, options) : null; +} + +/** + * Root that holds generated person skills (`-/SKILL.md`). + * Derived from the matrix so the two lists can never drift. + */ +export function generatedSkillsRoot(host, options = {}) { + const id = resolveHostId(host); + const override = GENERATED_ROOT_OVERRIDES[id]; + if (override) return expandTargetPath(override, options); + return dirname(repoInstallDir(id, options)); +} + +/** Python-compatible name for the generated-skill root (install_generated_skill.py). */ +export const defaultSkillsDir = generatedSkillsRoot; + +/** Refuse filesystem roots, the home directory and paths not named `distilly`. */ +export function validateInstallTarget(inputPath, { home = homedir(), requireName = true } = {}) { + const target = resolve(inputPath); + const parsed = parse(target); + if (target === parsed.root || target === resolve(home)) { + throw new Error("refusing to install into a filesystem root or home directory"); + } + if (requireName && basename(target) !== "distilly") { + throw new Error("the install path must end with a directory named distilly"); + } + return target; +} + +function pathsOverlap(source, destination) { + const sourceRoot = resolveRealPath(source); + const destinationRoot = resolveRealPath(destination); + if (sourceRoot === destinationRoot) return "same"; + const nested = + destinationRoot.startsWith(`${sourceRoot}/`) || sourceRoot.startsWith(`${destinationRoot}/`); + return nested ? "nested" : false; +} + +function shouldIgnore(name) { + return REPO_IGNORE.includes(name) || name.endsWith(".pyc"); +} + +/** Timestamped backup path used before replacing an existing install. */ +export function backupPathFor(target) { + const stamp = new Date().toISOString().replaceAll(":", "-").replaceAll(".", "-"); + return `${target}.backup-${stamp}`; +} + +/** + * Copy the Distilly repo into a host skill directory + * (`install_openclaw_skill.py` / `install_codex_skill.py` / `install_hermes_skill.py`). + */ +export function installRepoSkill({ + source, + destination, + force = false, + dryRun = false, + backup = false, +}) { + if (!existsSync(join(source, "SKILL.md"))) { + throw new Error(`source does not look like a skill repo: ${source}`); + } + + const overlap = pathsOverlap(source, destination); + if (overlap === "same") return destination; + if (overlap === "nested") throw new Error("source and destination must not overlap"); + + if (dryRun) return destination; + + let backupPath = null; + if (existsSync(destination)) { + if (!force) throw new Error(`destination already exists: ${destination}`); + if (backup) { + backupPath = backupPathFor(destination); + renameSync(destination, backupPath); + } else { + rmSync(destination, { recursive: true, force: true }); + } + } + + mkdirSync(dirname(destination), { recursive: true }); + cpSync(source, destination, { + recursive: true, + filter: (sourcePath) => !shouldIgnore(basename(sourcePath)), + }); + return { destination, backupPath }; +} + +/** + * Remove an installed Distilly copy. + * `--force` skips the "looks like a Distilly install" check; `--backup` keeps a + * timestamped copy instead of deleting. + */ +export function uninstallRepoSkill({ + destination, + force = false, + dryRun = false, + backup = false, + home = homedir(), +} = {}) { + const target = validateInstallTarget(destination, { home }); + if (!existsSync(target)) { + throw new Error(`nothing installed at ${target}`); + } + + const skillFile = join(target, "SKILL.md"); + if (!force) { + if (!existsSync(skillFile)) { + throw new Error(`${target} does not contain SKILL.md; rerun with --force to remove it anyway`); + } + const frontmatter = /^---\r?\n([\s\S]*?)\r?\n---/.exec(readFileSync(skillFile, "utf8")); + if (!frontmatter || !/^name:\s*distilly\s*$/m.test(frontmatter[1])) { + throw new Error(`${target} is not a Distilly install; rerun with --force to remove it anyway`); + } + } + + if (dryRun) return { destination: target, backupPath: null, removed: false }; + + if (backup) { + const backupPath = backupPathFor(target); + renameSync(target, backupPath); + return { destination: target, backupPath, removed: true }; + } + + rmSync(target, { recursive: true, force: true }); + return { destination: target, backupPath: null, removed: true }; +} + +const FRONTMATTER_RE = /^---\n([\s\S]*?)\n---\n?/; + +/** Load and normalize generated skill metadata from a skill directory. */ +export function loadGeneratedMeta(skillDir) { + const metaPath = join(skillDir, "meta.json"); + if (!existsSync(metaPath)) { + throw new Error(`generated skill is missing meta.json: ${skillDir}`); + } + return enrichExistingSkillMeta(JSON.parse(readFileSync(metaPath, "utf8")), skillDir); +} + +/** Rewrite the frontmatter name field to the installed command name. */ +export function rewriteFrontmatterName(markdown, newName) { + const match = FRONTMATTER_RE.exec(markdown); + if (!match) return markdown; + + const body = markdown.slice(match[0].length); + const lines = match[1].split(/\r?\n/); + const rewritten = []; + let replaced = false; + + for (const line of lines) { + if (line.startsWith("name:")) { + rewritten.push(`name: ${newName}`); + replaced = true; + } else { + rewritten.push(line); + } + } + if (!replaced) rewritten.unshift(`name: ${newName}`); + + return `---\n${rewritten.join("\n")}\n---\n\n${body.replace(/^\n+/, "")}`; +} + +/** Load a generated artifact and rewrite it for host installation. */ +export function renderInstalledMarkdown(skillDir, artifactName, commandName) { + const artifactPath = join(skillDir, artifactName); + if (!existsSync(artifactPath)) { + throw new Error(`generated artifact not found: ${artifactPath}`); + } + return rewriteFrontmatterName(readFileSync(artifactPath, "utf8"), commandName); +} + +/** Persist installation metadata for later debugging and upgrades. */ +export function writeInstallMetadata(installDir, payload) { + writeFileSync(join(installDir, ".distilly-install.json"), jsonDumps(payload), "utf8"); +} + +/** Windows installs also get a slash-command shim (install_claude_generated_skill.py). */ +export function shouldInstallCommandShim(systemName = process.platform) { + const current = String(systemName).toLowerCase(); + return current.startsWith("win"); +} + +/** + * Install a generated combined skill into a host skill directory + * (`install_generated_skill_common.py`). + */ +/** + * Directories a generated person Skill carries into a host install. + * + * The v2 layout puts the evidence *inside* the Skill so the host can read + * `knowledge/text`, the ledger, the derived claims and the rendered page while + * offline. Copying only `SKILL.md` (which is what this did) leaves every citation + * dangling at the destination — the page opens but every anchor points at nothing. + */ +export const CARRIED_DIRECTORIES = ["knowledge/raw", "knowledge/text", "evidence", "views", "assets"]; + +export function installGeneratedSkill({ + skillDir, + skillsDir, + force = false, + dryRun = false, + host, +}) { + const meta = loadGeneratedMeta(skillDir); + const artifacts = meta.artifacts; + const commandName = artifacts.combined_command; + const installedMarkdown = renderInstalledMarkdown( + skillDir, + artifacts.combined_skill, + commandName, + ); + + const installDir = join(skillsDir, commandName); + const installFile = join(installDir, "SKILL.md"); + + const overlap = pathsOverlap(skillDir, installDir); + if (overlap) { + throw new Error( + `generated skill source and install destination must not overlap: ${skillDir} -> ${installDir}`, + ); + } + + const installRecord = { + host, + command_name: commandName, + character: meta.character, + slug: meta.slug, + version: meta.version, + source_skill_dir: String(skillDir), + source_artifact: artifacts.combined_skill, + installed_at: nowIso(), + }; + + if (!dryRun) { + if (existsSync(installDir)) { + if (!force) throw new Error(`${host} skill already exists: ${installDir}`); + rmSync(installDir, { recursive: true, force: true }); + } + mkdirSync(installDir, { recursive: true }); + writeFileSync(installFile, installedMarkdown, "utf8"); + // Only directories that exist are copied, and the record lists what travelled, + // so "the host has the evidence" is checkable rather than assumed. + const carried = []; + for (const relativePath of CARRIED_DIRECTORIES) { + const source = join(skillDir, relativePath); + if (!existsSync(source)) continue; + cpSync(source, join(installDir, relativePath), { recursive: true }); + carried.push(relativePath); + } + const ledger = join(skillDir, "knowledge", "index.json"); + if (existsSync(ledger)) { + mkdirSync(join(installDir, "knowledge"), { recursive: true }); + cpSync(ledger, join(installDir, "knowledge", "index.json")); + carried.push("knowledge/index.json"); + } + if (carried.length > 0) installRecord.carried = carried; + writeInstallMetadata(installDir, installRecord); + } + + return { host, command_name: commandName, skill_dir: installDir, skill_file: installFile }; +} + +/** + * Claude Code variant: optional `~/.claude/commands/.md` shim + * (`install_claude_generated_skill.py`). + */ +export function installGeneratedSkillForClaude({ + skillDir, + skillsDir, + commandsDir = null, + force = false, + dryRun = false, + installCommandShim = false, +}) { + const result = installGeneratedSkill({ skillDir, skillsDir, force, dryRun, host: "claude-code" }); + const commandPath = commandsDir === null ? null : join(commandsDir, `${result.command_name}.md`); + + if (!dryRun && installCommandShim && commandPath !== null) { + const meta = loadGeneratedMeta(skillDir); + const installedMarkdown = renderInstalledMarkdown( + skillDir, + meta.artifacts.combined_skill, + result.command_name, + ); + mkdirSync(dirname(commandPath), { recursive: true }); + writeFileSync(commandPath, installedMarkdown, "utf8"); + } + + return { + ...result, + command_path: commandPath, + command_shim_installed: Boolean(installCommandShim && commandPath !== null), + }; +} + +/** Is a Distilly install present at this path? Used by `doctor`. */ +export function inspectInstall(target) { + const skillFile = join(target, "SKILL.md"); + if (!existsSync(target) || !existsSync(skillFile)) { + return { installed: false, path: target, version: null, bytes: 0 }; + } + const frontmatter = /^---\r?\n([\s\S]*?)\r?\n---/.exec(readFileSync(skillFile, "utf8")); + const version = frontmatter ? (/(?:^|\n)version:\s*"?([^"\n]+)"?/.exec(frontmatter[1])?.[1] ?? null) : null; + return { installed: true, path: target, version, bytes: statSync(skillFile).size }; +} From 8f0e67c9e8c0fd9ae6aaaccd089cd0e6a9111f36 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 08/15] fix(v2): print a registered command's bilingual help MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:fix(v2): print a registered command's bilingual help --- bin/distilly.mjs | 354 +++++++++++++++++++++------------------ src/commands/doctor.mjs | 159 ++++++++++++++++++ src/commands/index.mjs | 264 +++++++++++++++++++++++++++++ src/commands/install.mjs | 239 ++++++++++++++++++++++++++ src/commands/legacy.mjs | 243 +++++++++++++++++++++++++++ 5 files changed, 1093 insertions(+), 166 deletions(-) mode change 100755 => 100644 bin/distilly.mjs create mode 100644 src/commands/doctor.mjs create mode 100644 src/commands/index.mjs create mode 100644 src/commands/install.mjs create mode 100644 src/commands/legacy.mjs diff --git a/bin/distilly.mjs b/bin/distilly.mjs old mode 100755 new mode 100644 index 7460a406..f5330ab8 --- a/bin/distilly.mjs +++ b/bin/distilly.mjs @@ -1,204 +1,226 @@ #!/usr/bin/env node - -import { - cpSync, - existsSync, - mkdirSync, - readFileSync, - renameSync, - rmSync, -} from "node:fs"; -import { homedir } from "node:os"; -import { basename, dirname, join, parse, resolve } from "node:path"; +/** + * Distilly entry point. + * + * Contract: `docs/v2/CONTRACT.md` §1 — this is the only user-facing entry. It + * parses global flags, resolves a subcommand through the registry in + * `src/commands/index.mjs`, prints a bilingual help screen, and always answers + * with the receipt shape from §3 when `--json` is set. + * + * Adding a command: create `src/commands/.mjs`, call `register(...)` from + * it, and import that module below. See `docs/v2/NODE-CORE.md`. + */ + +import { existsSync, readFileSync } from "node:fs"; +import { join, resolve } from "node:path"; import { fileURLToPath } from "node:url"; -const packageRoot = fileURLToPath(new URL("..", import.meta.url)); -const packageMetadata = JSON.parse( - readFileSync(join(packageRoot, "package.json"), "utf8"), -); - -const payloadEntries = [ +import { ArgError, wantsHelp } from "../src/cli/args.mjs"; +import { isEntryPoint } from "../src/cli/entry.mjs"; +import { CliError, createReceipt, createReporter } from "../src/cli/receipt.mjs"; +// Importing the registry also registers every built-in command module. +import { + lookup, + missingCommandError, + renderCommandHelp, + renderHelp, + resolveCommand, +} from "../src/commands/index.mjs"; + +export const packageRoot = fileURLToPath(new URL("..", import.meta.url)); +const packageMetadata = JSON.parse(readFileSync(join(packageRoot, "package.json"), "utf8")); +const version = packageMetadata.version; +const binary = "distilly"; + +/** + * Every path the published package must contain. `src/` and `assets/` are not + * optional: this file imports `../src/cli/args.mjs` at startup and the viewer + * template is read from `assets/`. Kept in sync with `package.json`'s `files` + * by `validatePayload`, which refuses to pack when they disagree. + */ +export const payloadEntries = [ "SKILL.md", "prompts", "references", - "tools", - "requirements.txt", + "bin", + "src", + "assets", + "scripts", + "package.json", "INSTALL.md", "INSTALL_EN.md", "LICENSE", "CITATION.cff", ]; -const hosts = { - "claude-code": () => join(homedir(), ".claude", "skills", "distilly"), - openclaw: () => - join(homedir(), ".openclaw", "workspace", "skills", "distilly"), - hermes: () => - join(homedir(), ".hermes", "skills", "openclaw-imports", "distilly"), - codex: () => join(homedir(), ".agents", "skills", "distilly"), - "deepseek-harness": () => - join(process.env.DSH_HOME || join(homedir(), ".dsh"), "skills", "distilly"), - pi: () => join(homedir(), ".pi", "agent", "skills", "distilly"), - "grok-build": () => join(homedir(), ".grok", "skills", "distilly"), - opencode: () => - join(homedir(), ".config", "opencode", "skills", "distilly"), -}; - -const aliases = { - claude: "claude-code", - deepseek: "deepseek-harness", - grok: "grok-build", -}; - -function printHelp() { - console.log(`Distilly ${packageMetadata.version} - -Install the Distilly creator Skill into a supported agent host. - -Usage: - distilly install [--force] - distilly install --path [--force] - -Hosts: - claude-code, openclaw, hermes, codex, deepseek-harness, - pi, grok-build, opencode - -Options: - --force Preserve an existing install as a timestamped backup, then install - --path Install to a custom path whose final directory is named distilly - --version Print the package version - --help Show this help -`); -} - -function fail(message) { - console.error(`Error: ${message}`); - process.exit(1); +/** npm always includes these whatever `files` says, so they need no pattern. */ +const ALWAYS_PACKED = new Set(["package.json", "README.md", "LICENSE", "LICENCE"]); + +/** + * Does at least one `files` pattern put `entry` into the tarball? + * + * npm's `files` accepts bare names (`SKILL.md`), directories (`src/`) and globs + * (`prompts/**`); a directory pattern covers everything below it. Only the + * shapes this manifest actually uses are supported — an unrecognised pattern is + * reported as "does not cover", never silently treated as a match. + */ +function packs(entry, patterns) { + if (ALWAYS_PACKED.has(entry)) return true; + return patterns.some((raw) => { + const pattern = String(raw).replace(/\/+$/, ""); + if (pattern === entry) return true; + if (entry.startsWith(`${pattern}/`)) return true; + if (!pattern.includes("*")) return false; + const source = pattern + .split("**") + .map((part) => + part + .split("*") + .map((literal) => literal.replace(/[.+?^${}()|[\]\\]/g, "\\$&")) + .join("[^/]*"), + ) + .join(".*"); + return new RegExp(`^${source}$`).test(entry); + }); } -function validatePayload() { - const missing = payloadEntries.filter( - (entry) => !existsSync(join(packageRoot, entry)), - ); +/** + * Prepack guard: the published payload must be complete and version-consistent. + * + * This runs from `prepack`, so it must judge the tarball npm is about to build, + * not the working tree. Checking only `root` is what let a broken package ship: + * `files` still listed the Python-era `tools/` and `requirements.txt` and had + * dropped `src/` and `assets/`, `--check-package` printed "payload is valid" + * because the repo tree had everything, and the extracted tarball could not even + * start. So the manifest is checked too. + */ +export function validatePayload(root = packageRoot) { + const missing = payloadEntries.filter((entry) => !existsSync(join(root, entry))); if (missing.length > 0) { - fail(`package payload is missing: ${missing.join(", ")}`); + throw new CliError(`package payload is missing: ${missing.join(", ")}`, { + code: "payload-incomplete", + remedy: "restore the missing paths or update payloadEntries in bin/distilly.mjs.", + }); } - const skill = readFileSync(join(packageRoot, "SKILL.md"), "utf8"); - if (!skill.includes(`version: "${packageMetadata.version}"`)) { - fail("package.json version does not match SKILL.md"); + const manifest = JSON.parse(readFileSync(join(root, "package.json"), "utf8")); + const patterns = Array.isArray(manifest.files) ? manifest.files : []; + if (patterns.length === 0) { + throw new CliError("package.json declares no `files`, so the tarball would be unpredictable", { + code: "payload-manifest", + remedy: "add a `files` array listing bin/, src/, assets/, scripts/ and the documents.", + }); } -} -function expandHome(inputPath) { - if (inputPath === "~") return homedir(); - if (inputPath.startsWith("~/")) return join(homedir(), inputPath.slice(2)); - return inputPath; -} - -function validateTarget(inputPath) { - const target = resolve(expandHome(inputPath)); - const parsed = parse(target); - if (target === parsed.root || target === resolve(homedir())) { - fail("refusing to install into a filesystem root or home directory"); - } - if (basename(target) !== "distilly") { - fail("the install path must end with a directory named distilly"); + const absent = patterns + .map((raw) => String(raw).replace(/\/+$/, "")) + .filter((pattern) => !existsSync(join(root, pattern))); + if (absent.length > 0) { + throw new CliError(`package.json \`files\` names paths that do not exist: ${absent.join(", ")}`, { + code: "payload-manifest", + remedy: "remove the stale entries (or restore the paths) so the manifest describes this tree.", + }); } - return target; -} -function parseInstallArgs(args) { - let host; - let customPath; - let force = false; - - for (let index = 0; index < args.length; index += 1) { - const arg = args[index]; - if (arg === "--force") { - force = true; - } else if (arg === "--path") { - customPath = args[index + 1]; - if (!customPath) fail("--path requires a value"); - index += 1; - } else if (arg.startsWith("--")) { - fail(`unknown option: ${arg}`); - } else if (!host) { - host = aliases[arg] || arg; - } else { - fail(`unexpected argument: ${arg}`); - } + const omitted = payloadEntries.filter((entry) => !packs(entry, patterns)); + if (omitted.length > 0) { + throw new CliError(`package.json \`files\` would omit required paths: ${omitted.join(", ")}`, { + code: "payload-incomplete", + remedy: `add ${omitted.map((entry) => `"${entry}/"`).join(", ")} to \`files\`; the package cannot run without them.`, + }); } - if (customPath) return { target: validateTarget(customPath), force }; - if (!host) fail("choose a host or pass --path"); - if (!hosts[host]) fail(`unsupported host: ${host}`); - return { target: validateTarget(hosts[host]()), force }; + const skill = readFileSync(join(root, "SKILL.md"), "utf8"); + if (!skill.includes(`version: "${version}"`)) { + throw new CliError("package.json version does not match SKILL.md", { + code: "version-mismatch", + remedy: `set SKILL.md frontmatter version to "${version}" (or bump package.json).`, + }); + } } -function timestamp() { - return new Date().toISOString().replaceAll(":", "-").replaceAll(".", "-"); +function failureReceipt(command, error) { + return createReceipt(command, { + ok: false, + error: { + code: error.code ?? "error", + message: error.message, + ...(error.remedy ? { remedy: error.remedy } : {}), + }, + warnings: [error.message], + }); } -function install(target, force) { - validatePayload(); +async function main(argv) { + // `--json` is global (CONTRACT §1): every subcommand answers with a receipt. + const json = argv.includes("--json"); + const args = argv.filter((arg) => arg !== "--json"); + const reporter = createReporter(json); + + if (args.includes("--check-package")) { + validatePayload(); + // A validation diagnostic, not command output: `prepack` shares stdout with + // `npm pack --json`, which must stay parseable. + process.stderr.write("Distilly package payload is valid.\n"); + return 0; + } - if (existsSync(target) && !force) { - fail(`${target} already exists; rerun with --force to preserve and replace it`); + // Global flags only count before a command name: `skill version rollback + // --version v1` must reach the version manager, not print the CLI version. + if (args[0] === "--version") { + reporter.line(version); + return 0; } - const parent = dirname(target); - const staging = join(parent, `.distilly-install-${process.pid}`); - let backup; + const { name, rest } = resolveCommand(args); + const command = lookup(name); - mkdirSync(parent, { recursive: true }); - if (existsSync(staging)) { - fail(`temporary install path already exists: ${staging}`); + if (args.length === 0 || args[0] === "help" || (wantsHelp(args) && !command)) { + process.stdout.write(renderHelp({ version, binary })); + return 0; } - try { - mkdirSync(staging); - for (const entry of payloadEntries) { - cpSync(join(packageRoot, entry), join(staging, entry), { - recursive: true, - preserveTimestamps: true, - }); - } - - if (!existsSync(join(staging, "SKILL.md"))) { - throw new Error("staged install does not contain SKILL.md"); - } - - if (existsSync(target)) { - backup = `${target}.backup-${timestamp()}`; - renameSync(target, backup); - } - renameSync(staging, target); - } catch (error) { - if (existsSync(staging)) { - rmSync(staging, { recursive: true, force: true }); - } - if (backup && !existsSync(target) && existsSync(backup)) { - renameSync(backup, target); - } - throw error; + if (command === null) { + throw missingCommandError(name); } - console.log(`Distilly ${packageMetadata.version} installed at ${target}`); - if (backup) console.log(`Previous install preserved at ${backup}`); + if (wantsHelp(args)) { + process.stdout.write(`${renderCommandHelp(command, { binary })}\n`); + return 0; + } + + const result = (await command.run({ + argv: rest, + json, + reporter, + ctx: { packageRoot, version, binary }, + })) ?? {}; + + const receipt = result.receipt ?? createReceipt(command.name); + reporter.finish(receipt); + if (result.exitCode !== undefined) return result.exitCode; + return receipt.ok === false ? 1 : 0; } -const args = process.argv.slice(2); -if (args.includes("--check-package")) { - validatePayload(); - console.log("Distilly package payload is valid."); -} else if (args.includes("--version")) { - console.log(packageMetadata.version); -} else if (args.length === 0 || args.includes("--help") || args[0] === "help") { - printHelp(); -} else if (args[0] === "install") { - const { target, force } = parseInstallArgs(args.slice(1)); - install(target, force); -} else { - fail(`unknown command: ${args[0]}`); +// Dispatch only when this file is the process entry point. Importing it (tests do, +// to reach `validatePayload` and `payloadEntries`) must not run a command with the +// importer's argv. `isEntryPoint` resolves symlinks, so this still fires when the +// CLI is reached through an npm `bin` shim or any other symlinked path. +if (isEntryPoint(import.meta.url)) { + try { + process.exitCode = await main(process.argv.slice(2)); + } catch (error) { + const json = process.argv.includes("--json"); + const command = process.argv.slice(2).find((arg) => !arg.startsWith("-")) ?? null; + const reporter = createReporter(json); + const failure = + error instanceof CliError || error instanceof ArgError + ? error + : new CliError(error?.message ?? String(error), { code: "unexpected" }); + + reporter.warn(`Error: ${failure.message}`); + if (failure.remedy) reporter.warn(`Remedy: ${failure.remedy}`); + reporter.finish(failureReceipt(command, failure)); + process.exitCode = failure.exitCode ?? 1; + } } diff --git a/src/commands/doctor.mjs b/src/commands/doctor.mjs new file mode 100644 index 00000000..59c1c472 --- /dev/null +++ b/src/commands/doctor.mjs @@ -0,0 +1,159 @@ +/** + * `distilly doctor` — inventory health check (minimal, honest version). + * + * Checks what is actually on disk: + * 1. every host in the shared matrix: is Distilly installed there, at which + * version, and does the directory look like a Distilly install; + * 2. generated Skills under `skills//` (version, corrections); + * 3. ledger coverage whenever `knowledge/index.json` exists; + * 4. capabilities this build cannot run yet are listed in `unavailable` + * (CONTRACT §3: nothing is silently skipped). + */ + +import { existsSync, readFileSync, statSync } from "node:fs"; +import { join } from "node:path"; + +import { register, PLANNED } from "./index.mjs"; +import { createReceipt, describeFile, displayPath } from "../cli/receipt.mjs"; +import { parseArgs } from "../cli/args.mjs"; +import { listAgents } from "../hosts/agents.mjs"; +import { inspectInstall, repoInstallDir } from "../install/hosts.mjs"; +import { CHARACTER_PRESETS } from "../skill/presets.mjs"; +import { listSkills } from "../skill/writer.mjs"; + +function doctorHelp(binary = "distilly") { + const zh = [ + "用法:", + ` ${binary} doctor [--base-dir ] [--json]`, + "", + "检查项(最小版):", + " 1. 宿主矩阵里每个宿主是否已安装 Distilly(路径、SKILL.md 版本);", + " 2. 生成的 Skill 清单(skills//,含版本与 corrections);", + " 3. 账本覆盖率:有 knowledge/index.json 时统计条目与字节数;", + " 4. 未实现能力(parse/view/collect 等)写进回执的 unavailable,不静默跳过。", + ].join("\n"); + const en = [ + "Usage:", + ` ${binary} doctor [--base-dir ] [--json]`, + "", + "Checks (minimal):", + " 1. whether Distilly is installed for each host in the shared matrix (path, SKILL.md version);", + " 2. the generated Skill inventory (skills// with version and corrections);", + " 3. ledger coverage: entry and byte counts whenever knowledge/index.json exists;", + " 4. capabilities this build cannot run yet (parse/view/collect …) are reported in the receipt's unavailable list, never skipped silently.", + ].join("\n"); + return { zh, en }; +} + +const OPTIONS = { + "base-dir": { type: "string", value: "dir" }, +}; + +function readLedger(skillDir) { + const ledgerPath = join(skillDir, "knowledge", "index.json"); + if (!existsSync(ledgerPath)) { + return { path: ledgerPath, entries: 0, bytes: 0, anchors: 0, present: false }; + } + const text = readFileSync(ledgerPath, "utf8"); + let entries = []; + try { + const parsed = JSON.parse(text); + entries = Array.isArray(parsed) ? parsed : (parsed.entries ?? []); + } catch { + entries = []; + } + const anchors = entries.reduce((total, entry) => { + const list = entry?.anchors; + return total + (Array.isArray(list) ? list.length : 0); + }, 0); + return { + path: ledgerPath, + entries: entries.length, + bytes: statSync(ledgerPath).size, + anchors, + present: true, + }; +} + +register("doctor", { + summary: "体检宿主与 Skill 库存 / Health-check hosts and skill inventory", + usage: "distilly doctor [--base-dir ] [--json]", + options: OPTIONS, + ...doctorHelp(), + run({ argv, reporter }) { + const { flags } = parseArgs(argv, OPTIONS); + const warnings = []; + const inputs = []; + const outputs = []; + + reporter.line("Hosts / 宿主:"); + const hostRows = []; + for (const id of listAgents()) { + const target = repoInstallDir(id); + const state = inspectInstall(target); + hostRows.push({ host: id, path: displayPath(target), installed: state.installed, version: state.version }); + reporter.line( + ` ${state.installed ? "installed" : "missing "} ${id.padEnd(18)} ${displayPath(target)}${state.version ? ` (v${state.version})` : ""}`, + ); + if (state.installed) { + const described = describeFile(join(target, "SKILL.md")); + if (described) inputs.push(described); + } + } + + reporter.line(""); + reporter.line("Skills / 人物 Skill:"); + const familyBase = flags["base-dir"]; + let skillCount = 0; + let anchorTotal = 0; + for (const [family, preset] of Object.entries(CHARACTER_PRESETS)) { + if (preset.character !== family) continue; + const baseDir = familyBase ? join(familyBase, family) : (preset.storage_root ?? preset.legacy_storage_root); + const skills = listSkills(baseDir); + for (const skill of skills) { + skillCount += 1; + const skillDir = join(baseDir, skill.slug); + const ledger = readLedger(skillDir); + anchorTotal += ledger.anchors; + reporter.line( + ` ${family}/${skill.slug} ${skill.version} corrections=${skill.corrections_count} ` + + `knowledge=${ledger.present ? `${ledger.entries} entries / ${ledger.bytes} bytes` : "none"}`, + ); + if (ledger.present) { + const described = describeFile(ledger.path); + if (described) outputs.push(described); + } + const skillFile = describeFile(join(skillDir, "SKILL.md")); + if (skillFile) outputs.push(skillFile); + } + } + if (skillCount === 0) { + reporter.line(" none found (run `distilly skill create` first)"); + warnings.push("no generated skills found"); + } + + reporter.line(""); + reporter.line( + `Ledger coverage / 账本:${skillCount} skills, ${anchorTotal} anchors recorded, 0 cited ` + + "(evidence/derived is delivered by ds/06-retrospect)", + ); + + const unavailable = Object.entries(PLANNED).map(([command, branch]) => ({ + channel: command, + reason: `not implemented in this build; delivered by ${branch}`, + })); + reporter.line(""); + reporter.line(`Unavailable / 未实现:${unavailable.map((item) => item.channel).join(", ")}`); + + return { + receipt: createReceipt("doctor", { + inputs, + outputs, + anchors: { total: anchorTotal, cited: 0 }, + warnings, + unavailable, + }), + extra: { hosts: hostRows, skills: skillCount }, + }; + }, +}); diff --git a/src/commands/index.mjs b/src/commands/index.mjs new file mode 100644 index 00000000..bf67fd3b --- /dev/null +++ b/src/commands/index.mjs @@ -0,0 +1,264 @@ +/** + * Command registry — the single registration point for the Distilly CLI. + * + * `bin/distilly.mjs` only parses global flags and dispatches; every subcommand + * lives in its own module under `src/commands/` and registers itself here: + * + * ```js + * import { register } from "./index.mjs"; + * register("skill create", { + * summary: "创建一个 Skill / Create a Skill", + * usage: "distilly skill create [options]", + * options: {...}, // src/cli/args.mjs spec, for automatic usage + * run: async ({ flags, positionals, json, reporter, ctx }) => ({ receipt, lines }), + * }); + * ``` + * + * A definition returns `{receipt, lines, exitCode?}`; throwing `CliError` is the + * supported way to fail loudly with a remedy. See `docs/v2/NODE-CORE.md`. + * + * Names may contain one space (`view check`, `skill create`); dispatch prefers + * the two-token name. Commands promised by `docs/v2/CONTRACT.md` §1 but not yet + * implemented are listed in `PLANNED` together with the branch that owns them, + * so an unfinished build reports "not implemented" instead of "unknown command". + */ + +import { CliError } from "../cli/receipt.mjs"; + +/** + * Lazily created so the command modules can be imported at the bottom of this + * file: they call `register()` while this module body is still being evaluated, + * so the map must not depend on a top-level `const` having run yet. + */ +var REGISTRY; +function registry() { + if (!REGISTRY) REGISTRY = new Map(); + return REGISTRY; +} + +/** Commands frozen in CONTRACT §1 whose implementation ships in another branch. */ +export const PLANNED = { + // Contract §1 is fully implemented on this branch (the command surface, the + // eight collection channels, note, view, skill migrate). The map stays because + // `missingCommandError` and `doctor` read it to explain what is *not* here; + // it being empty is the signal that nothing is outstanding. +}; + +/** + * Register one subcommand. Duplicate names are a programming error and throw. + * @param {string} name + * @param {{summary: string, usage: string, options?: object, run: Function, hidden?: boolean}} definition + */ +export function register(name, definition) { + const map = registry(); + if (map.has(name)) throw new Error(`command already registered: ${name}`); + if (typeof definition?.run !== "function") { + throw new Error(`command ${name} needs a run() function`); + } + map.set(name, { name, hidden: false, options: {}, ...definition }); + return definition; +} + +export function lookup(name) { + return registry().get(name) ?? null; +} + +export function listCommandDetails({ includeHidden = false } = {}) { + return [...registry().values()] + .filter((command) => includeHidden || !command.hidden) + .sort((a, b) => a.name.localeCompare(b.name)); +} + +/** Registered command names (strings) — the shape other branches assert on. */ +export function listCommands({ includeHidden = false } = {}) { + return listCommandDetails({ includeHidden }).map((command) => command.name); +} + +/** + * Split argv into a command name and its remaining arguments. + * Two-token names win over one-token names (`view check` before `view`). + * + * Accepts either an argv array or a single command string — the acceptance + * script and several tests ask about a command by name ("parse-chat"), and a + * string is not an argv array: it must be split, not indexed per character. + * + * @param {string[]|string} tokens + */ +export function resolveCommand(tokens) { + const argv = typeof tokens === "string" ? tokens.trim().split(/\s+/).filter(Boolean) : [...tokens]; + if (argv.length >= 2) { + const twoToken = `${argv[0]} ${argv[1]}`; + if (registry().has(twoToken)) return { name: twoToken, rest: argv.slice(2) }; + } + if (argv.length >= 1 && registry().has(argv[0])) { + return { name: argv[0], rest: argv.slice(1) }; + } + return { name: argv[0] ?? null, rest: argv.slice(1) }; +} + +/** `null` when the command is registered, otherwise a loud, actionable error. */ +export function missingCommandError(name) { + const branch = PLANNED[name] ?? PLANNED[`${name} ${""}`.trim()]; + if (branch) { + return new CliError(`command not implemented in this build: ${name}`, { + code: "not-implemented", + remedy: `${name} is delivered by branch ${branch} (see docs/v2/STATUS.md); this branch (ds/01-node-core) ships skill/install/uninstall/doctor only.`, + }); + } + return new CliError(`unknown command: ${name}`, { + code: "unknown-command", + remedy: "run `distilly --help` for the command list.", + }); +} + +/** Terminal columns for one string (CJK counts as two), used for help tables. */ +export function displayWidth(text) { + let width = 0; + for (const character of text) { + const code = character.codePointAt(0); + const wide = + (code >= 0x1100 && code <= 0x115f) || + (code >= 0x2e80 && code <= 0xa4cf) || + (code >= 0xac00 && code <= 0xd7a3) || + (code >= 0xf900 && code <= 0xfaff) || + (code >= 0xfe30 && code <= 0xfe6f) || + (code >= 0xff00 && code <= 0xff60) || + (code >= 0xffe0 && code <= 0xffe6) || + (code >= 0x20000 && code <= 0x3fffd); + width += wide ? 2 : 1; + } + return width; +} + +/** Pad to `width` terminal columns so bilingual help tables line up. */ +export function padDisplay(text, width) { + const padding = Math.max(1, width - displayWidth(text)); + return `${text}${" ".repeat(padding)}`; +} + +const CATALOG_ZH = [ + ["skill create|update|list|version", "创建 / 更新 / 列出 / 归档 Skill(本分支)"], + ["install ", "把 Distilly 装进某个宿主的 skills 目录(本分支)"], + ["uninstall [|--path]", "卸载已安装的 Distilly(本分支)"], + ["doctor", "宿主与 Skill 库存体检(本分支,最小版)"], +]; + +const CATALOG_EN = [ + ["skill create|update|list|version", "create / update / list / archive Skills (this branch)"], + ["install ", "install Distilly into a host skills directory (this branch)"], + ["uninstall [|--path]", "remove an installed Distilly (this branch)"], + ["doctor", "host and skill inventory health check (this branch, minimal)"], +]; + +/** Bilingual help for the whole CLI (CONTRACT §6: 中文 → `---` → English). */ +export function renderHelp({ version, binary = "distilly" } = {}) { + const implemented = listCommandDetails(); + const plannedNames = Object.keys(PLANNED).sort(); + + const zh = [ + `Distilly ${version}`, + "", + "用法:", + ` ${binary} <命令> [选项]`, + ` ${binary} --help | --version`, + "", + "已实现:", + ...implemented.map((command) => ` ${command.usage.padEnd(46)}${command.summary.split(" / ")[0]}`), + "", + "契约中已冻结、由其他分支交付:", + ` ${plannedNames.join(", ")}`, + "", + "全局选项:", + " --json 以 JSON 回执输出(stdout 只有回执;人读信息走 stderr)", + " --help 显示帮助", + " --version 打印版本", + "", + "示例:", + ` ${binary} skill create --character colleague --name "Zadie Smith" --work work.md --persona persona.md`, + ` ${binary} skill list --character colleague`, + ` ${binary} install claude-code --json`, + ].join("\n"); + + const en = [ + `Distilly ${version}`, + "", + "Usage:", + ` ${binary} [options]`, + ` ${binary} --help | --version`, + "", + "Implemented:", + ...implemented.map((command) => ` ${command.usage.padEnd(46)}${command.summary.split(" / ")[1] ?? command.summary}`), + "", + "Frozen by the contract, delivered by other branches:", + ` ${plannedNames.join(", ")}`, + "", + "Global options:", + " --json emit a JSON receipt (stdout holds only the receipt; prose goes to stderr)", + " --help show this help", + " --version print the package version", + "", + "Examples:", + ` ${binary} skill create --character colleague --name "Zadie Smith" --work work.md --persona persona.md`, + ` ${binary} skill list --character colleague`, + ` ${binary} install claude-code --json`, + ].join("\n"); + + const catalogZh = ["命令总览:", ...CATALOG_ZH.map(([usage, text]) => ` ${usage.padEnd(46)}${text}`)].join("\n"); + const catalogEn = [ + "Command catalog:", + ...CATALOG_EN.map(([usage, text]) => ` ${usage.padEnd(46)}${text}`), + ].join("\n"); + + return `${zh}\n\n${catalogZh}\n\n---\n\n## English\n\n${en}\n\n${catalogEn}\n`; +} + +/** Usage string for one command, built from its registered options. */ +/** The bilingual pair, with the `---` divider and the `## English` marker the convention uses. */ +function bilingualBody(zh, en) { + const chinese = typeof zh === "string" && zh !== "" ? zh : ""; + const english = typeof en === "string" && en !== "" ? en : ""; + if (chinese === "") return english; + if (english === "") return chinese; + return `${chinese}\n\n---\n\n## English\n\n${english}`; +} + +export function renderCommandHelp(command, { binary = "distilly" } = {}) { + // Three shapes live in this tree, and all three must render: + // help: "…" a plain string + // help: { zh, en } the pair, kept under one key + // ...{ zh, en } the pair *spread*, which is what the registry convention + // actually produces — `command.help` is undefined here, + // which is why every --help used to print one line + let body = ""; + if (typeof command.help === "string") { + body = command.help; + } else if (command.help && typeof command.help === "object") { + body = bilingualBody(command.help.zh, command.help.en); + } else if (typeof command.zh === "string" || typeof command.en === "string") { + body = bilingualBody(command.zh, command.en); + } + const head = command.usage.replace(/^distilly/, binary); + return (body === "" ? head : `${head}\n\n${body}`).trimEnd(); +} + +/* ------------------------------------------------------------------ */ +/* built-in commands */ +/* ------------------------------------------------------------------ */ +/* Imported for their side effect: each module calls `register()`. They live at + the bottom because they import `register` from this file — the lazy REGISTRY + above is what makes that cycle safe. */ +import "./credentialed.mjs"; +import "./doctor.mjs"; +import "./harvest.mjs"; +import "./install.mjs"; +import "./legacy.mjs"; +import "./migrate.mjs"; +import "./note.mjs"; +import "./parse-chat.mjs"; +import "./parse-email.mjs"; +import "./parse-doc.mjs"; +import "./parse-archive.mjs"; +import "./parse-subtitle.mjs"; +import "./retrospect.mjs"; +import "./skill.mjs"; +import "./view.mjs"; diff --git a/src/commands/install.mjs b/src/commands/install.mjs new file mode 100644 index 00000000..a42a7488 --- /dev/null +++ b/src/commands/install.mjs @@ -0,0 +1,239 @@ +/** + * `distilly install ` / `distilly uninstall [|--path ]`. + * + * Host directories come from the shared matrix in `src/hosts/agents.mjs`; the + * copy / verify / remove actions live in `src/install/hosts.mjs` (the merged + * port of the eight `tools/install_*.py` scripts). + * + * `--force` replaces an existing install after renaming it to a timestamped + * backup (the pre-v2 CLI's safety behaviour); `--no-backup` deletes instead. + */ + +import { join } from "node:path"; + +import { register } from "./index.mjs"; +import { CliError, createReceipt, describeFile, displayPath, directoryBytes } from "../cli/receipt.mjs"; +import { parseArgs } from "../cli/args.mjs"; +import { listAgents } from "../hosts/agents.mjs"; +import { + HOST_ALIASES, + inspectInstall, + installRepoSkill, + repoInstallDir, + repoProjectDir, + resolveHostId, + supportedHosts, + uninstallRepoSkill, + validateInstallTarget, +} from "../install/hosts.mjs"; + +function installHelp(binary = "distilly") { + const zh = [ + "用法:", + ` ${binary} install [--force] [--dry-run] [--no-backup] [--project]`, + ` ${binary} install --path <以 distilly 结尾的目录> [--force]`, + "", + `宿主 (目录取自 src/hosts/agents.mjs,不猜路径):`, + ` ${supportedHosts().join(", ")}`, + `别名:${Object.entries(HOST_ALIASES).map(([alias, id]) => `${alias} → ${id}`).join(",")}`, + "", + "选项:", + " --force 覆盖已存在的安装(默认先把旧副本改名成带时间戳的备份)", + " --no-backup --force 时直接删除旧副本,不保留备份", + " --dry-run 只解析目标路径,不写盘", + " --project 装到项目级目录(仅当宿主有文档记载的项目目录)", + " --path 自定义安装目录(最后一级必须叫 distilly)", + "", + `卸载:${binary} uninstall |--path [--force] [--dry-run] [--backup]`, + ].join("\n"); + const en = [ + "Usage:", + ` ${binary} install [--force] [--dry-run] [--no-backup] [--project]`, + ` ${binary} install --path [--force]`, + "", + "Hosts (paths come from src/hosts/agents.mjs, nothing is guessed):", + ` ${supportedHosts().join(", ")}`, + `Aliases: ${Object.entries(HOST_ALIASES).map(([alias, id]) => `${alias} → ${id}`).join(", ")}`, + "", + "Options:", + " --force replace an existing install (the old copy is renamed to a timestamped backup first)", + " --no-backup with --force, delete the old copy instead of keeping a backup", + " --dry-run resolve the target path without writing", + " --project install into the project-local directory (only when the host documents one)", + " --path custom install directory whose final segment is distilly", + "", + `Uninstall: ${binary} uninstall |--path [--force] [--dry-run] [--backup]`, + ].join("\n"); + return { zh, en }; +} + +const INSTALL_OPTIONS = { + path: { type: "string", value: "dir" }, + force: { type: "boolean" }, + "no-backup": { type: "boolean" }, + "dry-run": { type: "boolean" }, + project: { type: "boolean" }, +}; + +const UNINSTALL_OPTIONS = { + path: { type: "string", value: "dir" }, + force: { type: "boolean" }, + "dry-run": { type: "boolean" }, + backup: { type: "boolean" }, + project: { type: "boolean" }, +}; + +/** Resolve the install target from `--path` or a host id. */ +export function resolveTarget(flags, { scope = "global" } = {}) { + if (flags.path) { + try { + return { target: validateInstallTarget(flags.path), host: null, scope }; + } catch (error) { + throw new CliError(error.message, { + code: "unsafe-target", + remedy: "choose a directory whose final segment is `distilly`.", + }); + } + } + if (!flags.host) { + throw new CliError("choose a host or pass --path", { + code: "usage", + remedy: `hosts: ${supportedHosts().join(", ")}`, + }); + } + let host; + try { + host = resolveHostId(flags.host); + } catch (error) { + throw new CliError(error.message, { + code: "unknown-host", + remedy: `hosts: ${supportedHosts().join(", ")}`, + }); + } + const useProject = Boolean(flags.project); + const target = useProject ? repoProjectDir(host) : repoInstallDir(host); + if (target === null) { + throw new CliError(`${host} has no documented project-local directory`, { + code: "no-project-path", + remedy: `install globally instead: distilly install ${host}`, + }); + } + return { target, host, scope: useProject ? "project" : "global" }; +} + +const installCommand = { + summary: "安装到宿主目录 / Install Distilly into a host skills directory", + usage: "distilly install > [--force] [--dry-run]", + options: INSTALL_OPTIONS, + ...installHelp(), + run({ argv, reporter, ctx }) { + const { flags, positionals } = parseArgs(argv, { ...INSTALL_OPTIONS, host: { type: "string" } }); + const host = positionals[0] ?? flags.host; + if (positionals.length > 1) { + throw new CliError(`unexpected argument: ${positionals[1]}`, { code: "usage" }); + } + const { target, scope } = resolveTarget({ ...flags, host }); + const backup = !flags["no-backup"]; + + const report = {}; + let destination = target; + try { + destination = installRepoSkill({ + source: ctx.packageRoot, + destination: target, + force: flags.force, + dryRun: flags["dry-run"], + backup, + report, + }); + } catch (error) { + throw new CliError(error.message, { + code: "install-failed", + remedy: flags.force + ? "check the target directory permissions." + : `rerun with --force to replace ${target} (a backup is kept unless --no-backup is passed).`, + }); + } + + if (flags["dry-run"]) { + reporter.line(`Would install Distilly ${ctx.version} at ${displayPath(destination)}`); + } else { + reporter.line(`Distilly ${ctx.version} installed at ${displayPath(destination)}`); + if (report.backupPath) { + reporter.line(`Previous install preserved at ${displayPath(report.backupPath)}`); + } + } + + const skillFile = describeFile(join(destination, "SKILL.md")); + return { + receipt: { + ...createReceipt("install", { + outputs: skillFile ? [skillFile] : [], + warnings: report.backupPath ? [`previous install preserved at ${displayPath(report.backupPath)}`] : [], + unavailable: listAgents() + .filter((id) => id !== resolveHostId(host ?? "")) + .map((id) => ({ channel: id, reason: "not the selected host" })), + }), + host: host ?? null, + scope, + dry_run: Boolean(flags["dry-run"]), + }, + }; + }, +}; + +const uninstallCommand = { + summary: "卸载 / Remove an installed Distilly from a host directory", + usage: "distilly uninstall [|--path ] [--force] [--dry-run]", + options: UNINSTALL_OPTIONS, + ...installHelp(), + run({ argv, reporter }) { + const { flags, positionals } = parseArgs(argv, { ...UNINSTALL_OPTIONS, host: { type: "string" } }); + const host = positionals[0] ?? flags.host; + const { target, scope } = resolveTarget({ ...flags, host }); + + let result; + try { + result = uninstallRepoSkill({ + destination: target, + force: flags.force, + dryRun: flags["dry-run"], + backup: flags.backup, + }); + } catch (error) { + throw new CliError(error.message, { + code: "uninstall-failed", + remedy: "pass --force only when you are sure this directory is a Distilly install.", + }); + } + + if (flags["dry-run"]) { + reporter.line(`Would remove ${displayPath(result.destination)}`); + } else if (result.backupPath) { + reporter.line(`Removed ${displayPath(result.destination)} (kept at ${displayPath(result.backupPath)})`); + } else { + reporter.line(`Removed ${displayPath(result.destination)}`); + } + + const before = inspectInstall(result.destination); + return { + receipt: { + ...createReceipt("uninstall", { + outputs: result.removed ? [{ path: displayPath(result.destination), bytes: directoryBytes(result.destination) }] : [], + warnings: result.backupPath ? [`kept a copy at ${displayPath(result.backupPath)}`] : [], + }), + host: host ?? null, + scope, + removed: result.removed, + was_installed: before.installed, + }, + }; + }, +}; + +export function registerInstallCommands() { + register("install", installCommand); + register("uninstall", uninstallCommand); +} + +registerInstallCommands(); diff --git a/src/commands/legacy.mjs b/src/commands/legacy.mjs new file mode 100644 index 00000000..8577dc37 --- /dev/null +++ b/src/commands/legacy.mjs @@ -0,0 +1,243 @@ +/** + * `distilly legacy [args...]` — migration adapter (CONTRACT §1). + * + * The pre-v2 command lines + * + * python3 tools/skill_writer.py --action create --slug x --name X … + * python3 tools/version_manager.py --action rollback --slug x --version v1 + * python3 tools/install_generated_skill.py --skill-dir … --host codex + * python3 tools/install_openclaw_skill.py --force + * + * are translated onto `skill create|update|list|version`, `install` and the + * merged installer, with a deprecation warning on stderr and the same exit code + * as the target command. It disappears again in PR③ together with the last + * Python entry point; `SKILL.md` is updated by ds/04-prompts. + */ + +import { existsSync } from "node:fs"; +import { homedir } from "node:os"; +import { join } from "node:path"; + +import { register, lookup } from "./index.mjs"; +import { CliError, createReceipt, describeFile, displayPath } from "../cli/receipt.mjs"; +import { ArgError, parseArgs } from "../cli/args.mjs"; +import { expandTargetPath, installGeneratedSkill, installGeneratedSkillForClaude } from "../install/hosts.mjs"; + +const WRITER_ACTIONS = { create: "skill create", update: "skill update", list: "skill list" }; +const VERSION_ACTIONS = ["list", "backup", "rollback", "cleanup"]; + +function deprecation(tool, target, reporter) { + reporter.warn( + `[deprecated] python3 tools/${tool} → ${target}(旧入口在 PR③ 移除 / legacy entry point is removed in PR③)`, + ); +} + +function runRegistered(name, argv, context, reporter) { + const command = lookup(name); + if (!command) { + throw new CliError(`legacy adapter target is not registered: ${name}`, { code: "not-implemented" }); + } + return command.run({ ...context, argv, reporter }); +} + +function splitAction(argv, tool, allowed) { + const flagIndexes = argv + .map((value, index) => (value === "--action" ? index : value.startsWith("--action=") ? index : -1)) + .filter((index) => index !== -1); + if (flagIndexes.length === 0) { + throw new CliError(`legacy ${tool} needs --action <${allowed.join("|")}>`, { code: "usage" }); + } + const index = flagIndexes[0]; + const raw = argv[index]; + const action = raw.includes("=") ? raw.slice(raw.indexOf("=") + 1) : argv[index + 1]; + if (!allowed.includes(action)) { + throw new CliError(`legacy ${tool}: unsupported --action ${action}`, { + code: "usage", + remedy: `supported: ${allowed.join(", ")}`, + }); + } + const rest = raw.includes("=") + ? [...argv.slice(0, index), ...argv.slice(index + 1)] + : [...argv.slice(0, index), ...argv.slice(index + 2)]; + return { action, rest }; +} + +function legacyWriter({ argv, json, reporter, ctx }) { + const { action, rest } = splitAction(argv, "skill_writer.py", Object.keys(WRITER_ACTIONS)); + const target = WRITER_ACTIONS[action]; + deprecation("skill_writer.py", `distilly ${target}`, reporter); + return runRegistered(target, rest, { json, reporter, ctx }, reporter); +} + +function legacyVersionManager({ argv, json, reporter, ctx }) { + const { action, rest } = splitAction(argv, "version_manager.py", VERSION_ACTIONS); + deprecation("version_manager.py", `distilly skill version ${action}`, reporter); + return runRegistered("skill version", [action, ...rest], { json, reporter, ctx }, reporter); +} + +const GENERATED_INSTALL_OPTIONS = { + "skill-dir": { type: "string", value: "dir" }, + host: { type: "string", value: "host" }, + "skills-dir": { type: "string", value: "dir" }, + "claude-skills-dir": { type: "string", value: "dir" }, + "claude-commands-dir": { type: "string", value: "dir" }, + "openclaw-skills-dir": { type: "string", value: "dir" }, + "codex-skills-dir": { type: "string", value: "dir" }, + "install-command-shim": { type: "boolean" }, + force: { type: "boolean" }, + "dry-run": { type: "boolean" }, +}; + +function legacyGeneratedInstaller(tool, host) { + return ({ argv, reporter }) => { + const { flags } = parseArgs(argv, GENERATED_INSTALL_OPTIONS); + if (!flags["skill-dir"]) { + throw new CliError(`legacy ${tool} needs --skill-dir`, { code: "usage" }); + } + const target = `distilly skill create --install-${host}-skill`; + deprecation(tool, target, reporter); + + const claude = host === "claude-code"; + const skillsDir = claude + ? flags["claude-skills-dir"] ?? join(homedir(), ".claude", "skills") + : flags["skills-dir"] ?? + flags[`${host}-skills-dir`] ?? + join(homedir(), host === "openclaw" ? ".openclaw/workspace/skills" : ".agents/skills"); + + const result = claude + ? installGeneratedSkillForClaude({ + skillDir: flags["skill-dir"], + skillsDir, + commandsDir: flags["claude-commands-dir"] ?? join(homedir(), ".claude", "commands"), + force: flags.force, + dryRun: flags["dry-run"], + installCommandShim: flags["install-command-shim"], + }) + : installGeneratedSkill({ + skillDir: flags["skill-dir"], + skillsDir, + force: flags.force, + dryRun: flags["dry-run"], + host, + }); + + reporter.line(result.command_name); + reporter.line(displayPath(result.skill_dir)); + if (result.command_shim_installed && result.command_path) { + reporter.line(displayPath(result.command_path)); + } + + const skillFile = describeFile(join(result.skill_dir, "SKILL.md")); + return { + receipt: createReceipt(`install ${host}`, { + outputs: skillFile ? [skillFile] : [], + warnings: [`deprecated entry point: tools/${tool}`], + }), + }; + }; +} + +const REPO_INSTALL_OPTIONS = { + source: { type: "string", value: "dir" }, + dest: { type: "string", value: "dir" }, + force: { type: "boolean" }, + "dry-run": { type: "boolean" }, +}; + +function legacyRepoInstaller(tool, host) { + return ({ argv, json, reporter, ctx }) => { + const { flags } = parseArgs(argv, REPO_INSTALL_OPTIONS); + deprecation(tool, `distilly install ${host}`, reporter); + const args = []; + if (flags.dest) args.push("--path", expandTargetPath(flags.dest)); + else args.push(host); + if (flags.force) args.push("--force"); + if (flags["dry-run"]) args.push("--dry-run"); + if (flags.source) { + throw new CliError(`legacy ${tool} --source is no longer configurable`, { + code: "unsupported-option", + remedy: "the CLI always installs the running checkout; run it from the source directory instead.", + }); + } + return runRegistered("install", args, { json, reporter, ctx }, reporter); + }; +} + +const LEGACY_TOOLS = { + "skill_writer.py": legacyWriter, + "version_manager.py": legacyVersionManager, + "install_generated_skill.py": legacyGeneratedInstaller("install_generated_skill.py", "codex"), + "install_claude_generated_skill.py": legacyGeneratedInstaller( + "install_claude_generated_skill.py", + "claude-code", + ), + "install_openclaw_generated_skill.py": legacyGeneratedInstaller( + "install_openclaw_generated_skill.py", + "openclaw", + ), + "install_codex_generated_skill.py": legacyGeneratedInstaller( + "install_codex_generated_skill.py", + "codex", + ), + "install_openclaw_skill.py": legacyRepoInstaller("install_openclaw_skill.py", "openclaw"), + "install_codex_skill.py": legacyRepoInstaller("install_codex_skill.py", "codex"), + "install_hermes_skill.py": legacyRepoInstaller("install_hermes_skill.py", "hermes"), +}; + +function legacyHelp(binary = "distilly") { + const tools = Object.keys(LEGACY_TOOLS).sort(); + const zh = [ + "用法(迁移期):", + ` ${binary} legacy [--action ...] [options]`, + "", + "支持的旧入口(转发到新子命令,并向 stderr 打印弃用警告):", + ...tools.map((tool) => ` tools/${tool}`), + "", + "示例:", + ` ${binary} legacy skill_writer.py --action create --slug eulalie --name Eulalie`, + ` ${binary} legacy version_manager.py --action rollback --slug eulalie --version v1`, + ].join("\n"); + const en = [ + "Usage (migration window):", + ` ${binary} legacy [--action ...] [options]`, + "", + "Supported legacy entry points (forwarded, with a deprecation warning on stderr):", + ...tools.map((tool) => ` tools/${tool}`), + "", + "Examples:", + ` ${binary} legacy skill_writer.py --action create --slug eulalie --name Eulalie`, + ` ${binary} legacy version_manager.py --action rollback --slug eulalie --version v1`, + ].join("\n"); + return { zh, en }; +} + +register("legacy", { + summary: "旧 python3 tools/*.py 入口转发 / Forward legacy python3 tools/*.py calls", + usage: "distilly legacy [--action ...] [options]", + hidden: true, + ...legacyHelp(), + run(context) { + const [tool, ...rest] = context.argv; + if (!tool) { + throw new CliError("legacy needs the tool name", { + code: "usage", + remedy: `supported: ${Object.keys(LEGACY_TOOLS).sort().join(", ")}`, + }); + } + const normalized = tool.replace(/^\.?\/?(tools\/)?/, ""); + const handler = LEGACY_TOOLS[normalized]; + if (!handler) { + throw new CliError(`no legacy adapter for ${tool}`, { + code: "unknown-tool", + remedy: + "collectors and research tools stay in Python until ds/07; use the new CLI for skill/install/uninstall/doctor.", + }); + } + if (!existsSync(join(context.ctx.packageRoot, "tools", normalized))) { + // The Python file is gone on this branch — forwarding is the whole point. + } + return handler({ ...context, argv: rest }); + }, +}); + +export { ArgError, LEGACY_TOOLS }; From 8360061b793aa1fa0dcfee96e576aa24cfc46dcf Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 09/15] feat(v2): derive the pinyin slug table from unihan MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:feat(v2): derive the pinyin slug table from unihan --- assets/pinyin.json | 3528 +++++++++++++++++++++++++++++++++++ scripts/generate-pinyin.mjs | 204 ++ tests/pinyin-slug.test.mjs | 131 ++ 3 files changed, 3863 insertions(+) create mode 100644 assets/pinyin.json create mode 100644 scripts/generate-pinyin.mjs create mode 100644 tests/pinyin-slug.test.mjs diff --git a/assets/pinyin.json b/assets/pinyin.json new file mode 100644 index 00000000..8793dbc5 --- /dev/null +++ b/assets/pinyin.json @@ -0,0 +1,3528 @@ +{ + "_comment": "Derived pinyin table for slug generation. Generated file — run `node scripts/generate-pinyin.mjs` to refresh; `--check` fails on drift.", + "source": { + "database": "Unihan", + "url": "https://www.unicode.org/Public/UCD/latest/ucd/Unihan.zip", + "file": "Unihan_Readings.txt", + "fields": [ + "kMandarin", + "kHanyuPinlu" + ], + "unicode_version": "17.0.0", + "source_date": "2025-07-24 00:00:00 GMT [KL]" + }, + "license": { + "name": "Unicode License v3", + "url": "https://www.unicode.org/license.txt", + "notice": "Unihan data is Copyright © Unicode, Inc. and distributed under the Unicode License v3; see the URL above for the full text." + }, + "generated_by": "scripts/generate-pinyin.mjs", + "selection": { + "rule": "top N Han characters by summed kHanyuPinlu frequency; reading = first kMandarin, else the most frequent kHanyuPinlu reading", + "limit": 3500, + "covered_characters": 44348 + }, + "count": 3500, + "characters": { + "的": "de", + "一": "yī", + "了": "le", + "是": "shì", + "不": "bù", + "我": "wǒ", + "在": "zài", + "有": "yǒu", + "人": "rén", + "这": "zhè", + "這": "zhè", + "他": "tā", + "们": "men", + "們": "men", + "來": "lái", + "来": "lái", + "个": "gè", + "個": "gè", + "上": "shàng", + "地": "de", + "大": "dà", + "著": "zhù", + "着": "zhe", + "就": "jiù", + "你": "nǐ", + "到": "dào", + "說": "shuō", + "说": "shuō", + "和": "hé", + "要": "yào", + "里": "lǐ", + "么": "me", + "子": "zǐ", + "也": "yě", + "去": "qù", + "那": "nà", + "得": "dé", + "主": "zhǔ", + "会": "huì", + "會": "huì", + "时": "shí", + "時": "shí", + "出": "chū", + "下": "xià", + "国": "guó", + "國": "guó", + "过": "guò", + "過": "guò", + "为": "wèi", + "為": "wèi", + "好": "hǎo", + "以": "yǐ", + "看": "kàn", + "可": "kě", + "还": "hái", + "還": "hái", + "生": "shēng", + "都": "dōu", + "学": "xué", + "學": "xué", + "沒": "méi", + "没": "méi", + "起": "qǐ", + "能": "néng", + "多": "duō", + "年": "nián", + "小": "xiǎo", + "把": "bǎ", + "天": "tiān", + "工": "gōng", + "家": "jiā", + "发": "fā", + "發": "fā", + "动": "dòng", + "動": "dòng", + "对": "duì", + "對": "duì", + "用": "yòng", + "中": "zhōng", + "又": "yòu", + "作": "zuò", + "同": "tóng", + "民": "mín", + "自": "zì", + "样": "yàng", + "樣": "yàng", + "想": "xiǎng", + "面": "miàn", + "成": "chéng", + "她": "tā", + "义": "yì", + "義": "yì", + "后": "hòu", + "經": "jīng", + "经": "jīng", + "产": "chǎn", + "產": "chǎn", + "十": "shí", + "什": "shén", + "道": "dào", + "进": "jìn", + "進": "jìn", + "现": "xiàn", + "現": "xiàn", + "只": "zhǐ", + "儿": "ér", + "兒": "ér", + "点": "diǎn", + "點": "diǎn", + "头": "tóu", + "頭": "tóu", + "种": "zhǒng", + "種": "zhǒng", + "从": "cóng", + "從": "cóng", + "甚": "shèn", + "些": "xiē", + "很": "hěn", + "然": "rán", + "前": "qián", + "見": "jiàn", + "见": "jiàn", + "老": "lǎo", + "事": "shì", + "方": "fāng", + "分": "fēn", + "于": "yú", + "开": "kāi", + "而": "ér", + "開": "kāi", + "麼": "me", + "心": "xīn", + "两": "liǎng", + "兩": "liǎng", + "走": "zǒu", + "行": "xíng", + "長": "zhǎng", + "长": "zhǎng", + "高": "gāo", + "象": "xiàng", + "三": "sān", + "当": "dāng", + "當": "dāng", + "它": "tā", + "氣": "qì", + "回": "huí", + "給": "gěi", + "给": "gěi", + "实": "shí", + "實": "shí", + "問": "wèn", + "问": "wèn", + "全": "quán", + "水": "shuǐ", + "部": "bù", + "几": "jǐ", + "二": "èr", + "命": "mìng", + "正": "zhèng", + "定": "dìng", + "党": "dǎng", + "黨": "dǎng", + "手": "shǒu", + "力": "lì", + "己": "jǐ", + "机": "jī", + "機": "jī", + "气": "qì", + "意": "yì", + "向": "xiàng", + "所": "suǒ", + "幾": "jǐ", + "知": "zhī", + "等": "děng", + "社": "shè", + "物": "wù", + "理": "lǐ", + "战": "zhàn", + "戰": "zhàn", + "边": "biān", + "邊": "biān", + "話": "huà", + "话": "huà", + "候": "hòu", + "但": "dàn", + "呢": "ne", + "声": "shēng", + "本": "běn", + "聲": "shēng", + "如": "rú", + "吶": "nà", + "呐": "nà", + "使": "shǐ", + "之": "zhī", + "打": "dǎ", + "叫": "jiào", + "外": "wài", + "罢": "bà", + "罷": "bà", + "法": "fǎ", + "眼": "yǎn", + "情": "qíng", + "做": "zuò", + "身": "shēn", + "重": "zhòng", + "化": "huà", + "革": "gé", + "才": "cái", + "間": "jiān", + "间": "jiān", + "反": "fǎn", + "已": "yǐ", + "四": "sì", + "最": "zuì", + "真": "zhēn", + "业": "yè", + "業": "yè", + "怎": "zěn", + "志": "zhì", + "听": "tīng", + "聽": "tīng", + "吧": "ba", + "別": "bié", + "别": "bié", + "級": "jí", + "级": "jí", + "放": "fàng", + "妈": "mā", + "媽": "mā", + "无": "wú", + "無": "wú", + "路": "lù", + "明": "míng", + "先": "xiān", + "干": "gàn", + "因": "yīn", + "新": "xīn", + "量": "liàng", + "車": "chē", + "车": "chē", + "文": "wén", + "阶": "jiē", + "階": "jiē", + "代": "dài", + "少": "shǎo", + "五": "wǔ", + "加": "jiā", + "解": "jiě", + "制": "zhì", + "政": "zhèng", + "军": "jūn", + "軍": "jūn", + "度": "dù", + "活": "huó", + "各": "gè", + "住": "zhù", + "电": "diàn", + "電": "diàn", + "比": "bǐ", + "员": "yuán", + "員": "yuán", + "第": "dì", + "常": "cháng", + "关": "guān", + "關": "guān", + "体": "tǐ", + "體": "tǐ", + "建": "jiàn", + "口": "kǒu", + "太": "tài", + "次": "cì", + "争": "zhēng", + "爭": "zhēng", + "月": "yuè", + "山": "shān", + "原": "yuán", + "再": "zài", + "吃": "chī", + "变": "biàn", + "變": "biàn", + "应": "yīng", + "應": "yīng", + "果": "guǒ", + "門": "mén", + "门": "mén", + "題": "tí", + "题": "tí", + "条": "tiáo", + "條": "tiáo", + "西": "xī", + "光": "guāng", + "思": "sī", + "由": "yóu", + "快": "kuài", + "利": "lì", + "表": "biǎo", + "东": "dōng", + "東": "dōng", + "总": "zǒng", + "總": "zǒng", + "您": "nín", + "合": "hé", + "立": "lì", + "百": "bǎi", + "提": "tí", + "吗": "ma", + "嗎": "ma", + "被": "bèi", + "跟": "gēn", + "領": "lǐng", + "领": "lǐng", + "結": "jié", + "结": "jié", + "啊": "a", + "决": "jué", + "決": "jué", + "完": "wán", + "平": "píng", + "教": "jiào", + "队": "duì", + "隊": "duì", + "論": "lùn", + "论": "lùn", + "許": "xǔ", + "许": "xǔ", + "科": "kē", + "其": "qí", + "亲": "qīn", + "親": "qīn", + "資": "zī", + "资": "zī", + "者": "zhě", + "九": "jiǔ", + "展": "zhǎn", + "书": "shū", + "書": "shū", + "內": "nèi", + "内": "nèi", + "更": "gèng", + "并": "bìng", + "呀": "ya", + "哪": "nǎ", + "导": "dǎo", + "導": "dǎo", + "笑": "xiào", + "性": "xìng", + "白": "bái", + "系": "xì", + "造": "zào", + "斗": "dòu", + "相": "xiāng", + "带": "dài", + "帶": "dài", + "万": "wàn", + "萬": "wàn", + "敌": "dí", + "敵": "dí", + "指": "zhǐ", + "界": "jiè", + "共": "gòng", + "接": "jiē", + "直": "zhí", + "便": "biàn", + "公": "gōng", + "往": "wǎng", + "农": "nóng", + "農": "nóng", + "線": "xiàn", + "线": "xiàn", + "記": "jì", + "记": "jì", + "日": "rì", + "位": "wèi", + "認": "rèn", + "认": "rèn", + "每": "měi", + "研": "yán", + "今": "jīn", + "世": "shì", + "将": "jiāng", + "將": "jiāng", + "任": "rèn", + "孩": "hái", + "根": "gēn", + "花": "huā", + "难": "nán", + "難": "nán", + "区": "qū", + "區": "qū", + "覺": "jué", + "觉": "jué", + "群": "qún", + "运": "yùn", + "運": "yùn", + "办": "bàn", + "辦": "bàn", + "風": "fēng", + "风": "fēng", + "数": "shù", + "數": "shù", + "望": "wàng", + "究": "jiū", + "識": "shí", + "识": "shí", + "写": "xiě", + "寫": "xiě", + "处": "chù", + "處": "chù", + "女": "nǚ", + "治": "zhì", + "件": "jiàn", + "流": "liú", + "却": "què", + "卻": "què", + "众": "zhòng", + "眾": "zhòng", + "半": "bàn", + "师": "shī", + "師": "shī", + "通": "tōng", + "愛": "ài", + "爱": "ài", + "或": "huò", + "拿": "ná", + "八": "bā", + "形": "xíng", + "步": "bù", + "此": "cǐ", + "計": "jì", + "计": "jì", + "必": "bì", + "站": "zhàn", + "特": "tè", + "設": "shè", + "设": "shè", + "改": "gǎi", + "受": "shòu", + "连": "lián", + "連": "lián", + "信": "xìn", + "切": "qiè", + "誰": "shuí", + "谁": "shéi", + "強": "qiáng", + "强": "qiáng", + "該": "gāi", + "该": "gāi", + "朮": "shù", + "术": "shù", + "且": "qiě", + "找": "zhǎo", + "算": "suàn", + "远": "yuǎn", + "遠": "yuǎn", + "六": "liù", + "满": "mǎn", + "滿": "mǎn", + "觀": "guān", + "观": "guān", + "早": "zǎo", + "報": "bào", + "报": "bào", + "坐": "zuò", + "热": "rè", + "熱": "rè", + "期": "qī", + "济": "jì", + "濟": "jì", + "石": "shí", + "似": "sì", + "告": "gào", + "够": "gòu", + "夠": "gòu", + "跑": "pǎo", + "啦": "la", + "管": "guǎn", + "料": "liào", + "感": "gǎn", + "爸": "bà", + "講": "jiǎng", + "讓": "ràng", + "让": "ràng", + "讲": "jiǎng", + "与": "yǔ", + "與": "yǔ", + "組": "zǔ", + "组": "zǔ", + "統": "tǒng", + "统": "tǒng", + "河": "hé", + "越": "yuè", + "火": "huǒ", + "爷": "yé", + "爺": "yé", + "服": "fú", + "七": "qī", + "色": "sè", + "飛": "fēi", + "飞": "fēi", + "轉": "zhuǎn", + "转": "zhuǎn", + "死": "sǐ", + "脸": "liǎn", + "臉": "liǎn", + "块": "kuài", + "塊": "kuài", + "确": "què", + "確": "què", + "空": "kōng", + "船": "chuán", + "务": "wù", + "務": "wù", + "取": "qǔ", + "场": "chǎng", + "場": "chǎng", + "海": "hǎi", + "极": "jí", + "極": "jí", + "質": "zhì", + "质": "zhì", + "准": "zhǔn", + "紧": "jǐn", + "緊": "jǐn", + "整": "zhěng", + "倒": "dào", + "基": "jī", + "錢": "qián", + "钱": "qián", + "馬": "mǎ", + "马": "mǎ", + "团": "tuán", + "團": "tuán", + "照": "zhào", + "千": "qiān", + "品": "pǐn", + "神": "shén", + "刚": "gāng", + "剛": "gāng", + "怕": "pà", + "輕": "qīng", + "轻": "qīng", + "土": "tǔ", + "劳": "láo", + "勞": "láo", + "树": "shù", + "樹": "shù", + "影": "yǐng", + "保": "bǎo", + "史": "shǐ", + "細": "xì", + "细": "xì", + "紅": "hóng", + "红": "hóng", + "习": "xí", + "習": "xí", + "程": "chéng", + "青": "qīng", + "近": "jìn", + "容": "róng", + "油": "yóu", + "历": "lì", + "歷": "lì", + "清": "qīng", + "求": "qiú", + "送": "sòng", + "錯": "cuò", + "错": "cuò", + "字": "zì", + "目": "mù", + "村": "cūn", + "裡": "lǐ", + "据": "jù", + "據": "jù", + "席": "xí", + "片": "piàn", + "夜": "yè", + "較": "jiào", + "较": "jiào", + "响": "xiǎng", + "響": "xiǎng", + "类": "lèi", + "類": "lèi", + "驗": "yàn", + "验": "yàn", + "离": "lí", + "離": "lí", + "底": "dǐ", + "至": "zhì", + "张": "zhāng", + "張": "zhāng", + "備": "bèi", + "入": "rù", + "备": "bèi", + "米": "mǐ", + "买": "mǎi", + "屋": "wū", + "買": "mǎi", + "深": "shēn", + "器": "qì", + "收": "shōu", + "名": "míng", + "咱": "zán", + "規": "guī", + "规": "guī", + "集": "jí", + "需": "xū", + "南": "nán", + "勝": "shèng", + "胜": "shèng", + "布": "bù", + "病": "bìng", + "具": "jù", + "鐵": "tiě", + "铁": "tiě", + "須": "xū", + "须": "xū", + "織": "zhī", + "织": "zhī", + "装": "zhuāng", + "裝": "zhuāng", + "厂": "chǎng", + "廠": "chǎng", + "晚": "wǎn", + "北": "běi", + "睛": "jīng", + "及": "jí", + "况": "kuàng", + "況": "kuàng", + "院": "yuàn", + "传": "chuán", + "傳": "chuán", + "友": "yǒu", + "技": "jì", + "哥": "gē", + "房": "fáng", + "消": "xiāo", + "包": "bāo", + "际": "jì", + "際": "jì", + "母": "mǔ", + "坚": "jiān", + "堅": "jiān", + "批": "pī", + "談": "tán", + "谈": "tán", + "何": "hé", + "市": "shì", + "黑": "hēi", + "非": "fēi", + "忙": "máng", + "断": "duàn", + "斷": "duàn", + "赶": "gǎn", + "趕": "gǎn", + "汽": "qì", + "族": "zú", + "睡": "shuì", + "拉": "lā", + "委": "wěi", + "速": "sù", + "低": "dī", + "精": "jīng", + "兴": "xìng", + "抗": "kàng", + "興": "xìng", + "害": "hài", + "围": "wéi", + "圍": "wéi", + "刻": "kè", + "派": "pài", + "答": "dá", + "衣": "yī", + "苦": "kǔ", + "击": "jī", + "擊": "jī", + "交": "jiāo", + "娘": "niáng", + "支": "zhī", + "音": "yīn", + "严": "yán", + "嚴": "yán", + "广": "guǎng", + "廣": "guǎng", + "脚": "jiǎo", + "腳": "jiǎo", + "压": "yā", + "壓": "yā", + "句": "jù", + "急": "jí", + "坏": "huài", + "壞": "huài", + "草": "cǎo", + "嘴": "zuǐ", + "艺": "yì", + "藝": "yì", + "始": "shǐ", + "帝": "dì", + "破": "pò", + "单": "dān", + "單": "dān", + "調": "diào", + "调": "diào", + "专": "zhuān", + "專": "zhuān", + "增": "zēng", + "持": "chí", + "随": "suí", + "隨": "suí", + "帮": "bāng", + "幫": "bāng", + "安": "ān", + "訴": "sù", + "诉": "sù", + "穿": "chuān", + "城": "chéng", + "乎": "hū", + "士": "shì", + "請": "qǐng", + "请": "qǐng", + "联": "lián", + "聯": "lián", + "式": "shì", + "阵": "zhèn", + "陣": "zhèn", + "伟": "wěi", + "偉": "wěi", + "議": "yì", + "议": "yì", + "客": "kè", + "金": "jīn", + "星": "xīng", + "般": "bān", + "积": "jī", + "積": "jī", + "商": "shāng", + "复": "fù", + "複": "fù", + "达": "dá", + "達": "dá", + "飯": "fàn", + "饭": "fàn", + "約": "yuē", + "约": "yuē", + "虫": "chóng", + "参": "cān", + "參": "cān", + "举": "jǔ", + "亮": "liàng", + "舉": "jǔ", + "桥": "qiáo", + "橋": "qiáo", + "育": "yù", + "左": "zuǒ", + "雨": "yǔ", + "虽": "suī", + "雖": "suī", + "魚": "yú", + "鱼": "yú", + "兵": "bīng", + "毛": "máo", + "则": "zé", + "則": "zé", + "忽": "hū", + "節": "jié", + "节": "jié", + "推": "tuī", + "段": "duàn", + "卖": "mài", + "台": "tái", + "易": "yì", + "賣": "mài", + "落": "luò", + "鋼": "gāng", + "钢": "gāng", + "失": "shī", + "愿": "yuàn", + "材": "cái", + "靠": "kào", + "伙": "huǒ", + "岁": "suì", + "歲": "suì", + "皮": "pí", + "証": "zhèng", + "证": "zhèng", + "父": "fù", + "朋": "péng", + "阳": "yáng", + "陽": "yáng", + "即": "jí", + "微": "wēi", + "誤": "wù", + "误": "wù", + "停": "tíng", + "示": "shì", + "划": "huà", + "局": "jú", + "背": "bèi", + "显": "xiǎn", + "顯": "xiǎn", + "欢": "huān", + "歡": "huān", + "夫": "fū", + "引": "yǐn", + "息": "xī", + "除": "chú", + "温": "wēn", + "溫": "wēn", + "画": "huà", + "畫": "huà", + "食": "shí", + "首": "shǒu", + "图": "tú", + "圖": "tú", + "右": "yòu", + "号": "hào", + "號": "hào", + "續": "xù", + "续": "xù", + "层": "céng", + "層": "céng", + "呼": "hū", + "留": "liú", + "敢": "gǎn", + "权": "quán", + "權": "quán", + "灯": "dēng", + "燈": "dēng", + "密": "mì", + "旧": "jiù", + "舊": "jiù", + "静": "jìng", + "靜": "jìng", + "另": "lìng", + "突": "tū", + "掉": "diào", + "旁": "páng", + "查": "chá", + "跳": "tiào", + "护": "hù", + "護": "hù", + "久": "jiǔ", + "紀": "jì", + "纪": "jì", + "紙": "zhǐ", + "纸": "zhǐ", + "美": "měi", + "雪": "xuě", + "修": "xiū", + "助": "zhù", + "喊": "hǎn", + "冲": "chōng", + "沖": "chōng", + "医": "yī", + "存": "cún", + "醫": "yī", + "喜": "xǐ", + "渐": "jiàn", + "漸": "jiàn", + "球": "qiú", + "姑": "gū", + "呵": "hē", + "激": "jī", + "令": "lìng", + "冷": "lěng", + "势": "shì", + "勢": "shì", + "创": "chuàng", + "創": "chuàng", + "弟": "dì", + "念": "niàn", + "沉": "chén", + "注": "zhù", + "略": "lüè", + "頂": "dǐng", + "顶": "dǐng", + "古": "gǔ", + "律": "lǜ", + "按": "àn", + "評": "píng", + "评": "píng", + "脑": "nǎo", + "腦": "nǎo", + "室": "shì", + "搞": "gǎo", + "乡": "xiāng", + "唱": "chàng", + "鄉": "xiāng", + "府": "fǔ", + "讀": "dú", + "读": "dú", + "简": "jiǎn", + "簡": "jiǎn", + "价": "jià", + "價": "jià", + "养": "yǎng", + "板": "bǎn", + "養": "yǎng", + "县": "xiàn", + "縣": "xiàn", + "校": "xiào", + "烈": "liè", + "惊": "jīng", + "沙": "shā", + "驚": "jīng", + "章": "zhāng", + "視": "shì", + "视": "shì", + "采": "cǎi", + "維": "wéi", + "维": "wéi", + "血": "xuè", + "姐": "jiě", + "慢": "màn", + "故": "gù", + "木": "mù", + "怪": "guài", + "鐘": "zhōng", + "钟": "zhōng", + "省": "shěng", + "药": "yào", + "藥": "yào", + "角": "jiǎo", + "初": "chū", + "繼": "jì", + "继": "jì", + "抓": "zhuā", + "班": "bān", + "仅": "jǐn", + "僅": "jǐn", + "排": "pái", + "奶": "nǎi", + "封": "fēng", + "础": "chǔ", + "礎": "chǔ", + "烧": "shāo", + "燒": "shāo", + "周": "zhōu", + "喝": "hē", + "座": "zuò", + "担": "dān", + "擔": "dān", + "伤": "shāng", + "傷": "shāng", + "央": "yāng", + "棉": "mián", + "竟": "jìng", + "搖": "yáo", + "摇": "yáo", + "曾": "céng", + "困": "kùn", + "枪": "qiāng", + "槍": "qiāng", + "熟": "shú", + "終": "zhōng", + "终": "zhōng", + "功": "gōng", + "态": "tài", + "態": "tài", + "止": "zhǐ", + "源": "yuán", + "床": "chuáng", + "仍": "réng", + "尽": "jǐn", + "盡": "jǐn", + "懂": "dǒng", + "弹": "dàn", + "彈": "dàn", + "充": "chōng", + "防": "fáng", + "試": "shì", + "试": "shì", + "双": "shuāng", + "哭": "kū", + "雙": "shuāng", + "窗": "chuāng", + "吸": "xī", + "例": "lì", + "属": "shǔ", + "屬": "shǔ", + "翻": "fān", + "叔": "shū", + "祖": "zǔ", + "挥": "huī", + "揮": "huī", + "游": "yóu", + "缺": "quē", + "責": "zé", + "责": "zé", + "模": "mó", + "野": "yě", + "乱": "luàn", + "亂": "luàn", + "杂": "zá", + "痛": "tòng", + "适": "shì", + "適": "shì", + "雜": "zá", + "歌": "gē", + "菜": "cài", + "替": "tì", + "换": "huàn", + "換": "huàn", + "妇": "fù", + "婦": "fù", + "烟": "yān", + "煙": "yān", + "負": "fù", + "负": "fù", + "黃": "huáng", + "黄": "huáng", + "奇": "qí", + "瞭": "liào", + "占": "zhàn", + "岸": "àn", + "标": "biāo", + "標": "biāo", + "待": "dài", + "依": "yī", + "侵": "qīn", + "值": "zhí", + "林": "lín", + "課": "kè", + "课": "kè", + "卫": "wèi", + "衛": "wèi", + "嘛": "ma", + "选": "xuǎn", + "選": "xuǎn", + "称": "chēng", + "稱": "chēng", + "乐": "lè", + "庄": "zhuāng", + "握": "wò", + "检": "jiǎn", + "樂": "lè", + "檢": "jiǎn", + "武": "wǔ", + "莊": "zhuāng", + "田": "tián", + "益": "yì", + "街": "jiē", + "嫂": "sǎo", + "考": "kǎo", + "巨": "jù", + "演": "yǎn", + "營": "yíng", + "营": "yíng", + "爬": "pá", + "暗": "àn", + "未": "wèi", + "滅": "miè", + "灭": "miè", + "貨": "huò", + "货": "huò", + "差": "chà", + "春": "chūn", + "固": "gù", + "元": "yuán", + "顧": "gù", + "顾": "gù", + "普": "pǔ", + "希": "xī", + "含": "hán", + "弄": "nòng", + "針": "zhēn", + "针": "zhēn", + "短": "duǎn", + "降": "jiàng", + "型": "xíng", + "斤": "jīn", + "构": "gòu", + "架": "jià", + "格": "gé", + "構": "gòu", + "供": "gōng", + "透": "tòu", + "射": "shè", + "富": "fù", + "致": "zhì", + "副": "fù", + "攻": "gōng", + "忘": "wàng", + "践": "jiàn", + "踐": "jiàn", + "足": "zú", + "司": "sī", + "危": "wēi", + "既": "jì", + "泥": "ní", + "笔": "bǐ", + "筆": "bǐ", + "伸": "shēn", + "言": "yán", + "朝": "cháo", + "迫": "pò", + "抬": "tái", + "費": "fèi", + "费": "fèi", + "景": "jǐng", + "永": "yǒng", + "哎": "āi", + "叶": "yè", + "葉": "yè", + "减": "jiǎn", + "印": "yìn", + "店": "diàn", + "減": "jiǎn", + "江": "jiāng", + "宣": "xuān", + "洋": "yáng", + "劲": "jìn", + "勁": "jìn", + "某": "mǒu", + "絕": "jué", + "绝": "jué", + "抱": "bào", + "掌": "zhǎng", + "环": "huán", + "環": "huán", + "配": "pèi", + "遍": "biàn", + "映": "yìng", + "素": "sù", + "謝": "xiè", + "谢": "xiè", + "互": "hù", + "嗯": "ǹg", + "察": "chá", + "洗": "xǐ", + "优": "yōu", + "余": "yú", + "優": "yōu", + "概": "gài", + "桌": "zhuō", + "鼓": "gǔ", + "鏡": "jìng", + "镜": "jìng", + "刀": "dāo", + "摸": "mō", + "效": "xiào", + "味": "wèi", + "奋": "fèn", + "奮": "fèn", + "怀": "huái", + "懷": "huái", + "唯": "wéi", + "境": "jìng", + "粮": "liáng", + "糧": "liáng", + "肯": "kěn", + "楚": "chǔ", + "盾": "dùn", + "矛": "máo", + "王": "wáng", + "肉": "ròu", + "討": "tǎo", + "讨": "tǎo", + "官": "guān", + "摆": "bǎi", + "擺": "bǎi", + "杀": "shā", + "殺": "shā", + "逐": "zhú", + "筑": "zhù", + "仿": "fǎng", + "燃": "rán", + "冬": "dōng", + "袋": "dài", + "追": "zhuī", + "列": "liè", + "午": "wǔ", + "宝": "bǎo", + "寶": "bǎo", + "挂": "guà", + "掛": "guà", + "牛": "niú", + "置": "zhì", + "状": "zhuàng", + "狀": "zhuàng", + "鞋": "xié", + "假": "jiǎ", + "順": "shùn", + "顺": "shùn", + "丰": "fēng", + "墙": "qiáng", + "投": "tóu", + "牆": "qiáng", + "独": "dú", + "獨": "dú", + "矿": "kuàng", + "礦": "kuàng", + "腿": "tuǐ", + "酒": "jiǔ", + "語": "yǔ", + "语": "yǔ", + "遇": "yù", + "哦": "ó", + "浪": "làng", + "端": "duān", + "策": "cè", + "园": "yuán", + "園": "yuán", + "妹": "mèi", + "猛": "měng", + "幸": "xìng", + "彻": "chè", + "徹": "chè", + "炼": "liàn", + "煉": "liàn", + "碎": "suì", + "超": "chāo", + "案": "àn", + "退": "tuì", + "闹": "nào", + "鬧": "nào", + "佛": "fú", + "判": "pàn", + "英": "yīng", + "努": "nǔ", + "閃": "shǎn", + "闪": "shǎn", + "煤": "méi", + "犯": "fàn", + "瞧": "qiáo", + "散": "sàn", + "男": "nán", + "湖": "hú", + "鮮": "xiān", + "鲜": "xiān", + "骨": "gǔ", + "枝": "zhī", + "練": "liàn", + "练": "liàn", + "企": "qǐ", + "抽": "chōu", + "雞": "jī", + "鸡": "jī", + "銀": "yín", + "银": "yín", + "朵": "duǒ", + "露": "lù", + "館": "guǎn", + "馆": "guǎn", + "限": "xiàn", + "吹": "chuī", + "挺": "tǐng", + "脫": "tuō", + "脱": "tuō", + "婶": "shěn", + "嬸": "shěn", + "季": "jì", + "洞": "dòng", + "盖": "gài", + "碗": "wǎn", + "蓋": "gài", + "項": "xiàng", + "项": "xiàng", + "召": "zhào", + "像": "xiàng", + "堆": "duī", + "泪": "lèi", + "淚": "lèi", + "姓": "xìng", + "折": "zhé", + "束": "shù", + "沿": "yán", + "率": "lǜ", + "輸": "shū", + "输": "shū", + "否": "fǒu", + "哈": "hā", + "削": "xuē", + "套": "tào", + "汉": "hàn", + "测": "cè", + "測": "cè", + "漢": "hàn", + "哩": "lī", + "毫": "háo", + "鬼": "guǐ", + "勇": "yǒng", + "拍": "pāi", + "玩": "wán", + "輪": "lún", + "轮": "lún", + "险": "xiǎn", + "險": "xiǎn", + "巴": "bā", + "硬": "yìng", + "移": "yí", + "耐": "nài", + "震": "zhèn", + "預": "yù", + "预": "yù", + "临": "lín", + "綠": "lǜ", + "绿": "lǜ", + "股": "gǔ", + "臨": "lín", + "倍": "bèi", + "碰": "pèng", + "執": "zhí", + "守": "shǒu", + "悄": "qiāo", + "执": "zhí", + "敗": "bài", + "染": "rǎn", + "植": "zhí", + "败": "bài", + "鋪": "pù", + "铺": "pù", + "哲": "zhé", + "戶": "hù", + "户": "hù", + "伍": "wǔ", + "救": "jiù", + "狗": "gǒu", + "羊": "yáng", + "鎮": "zhèn", + "镇": "zhèn", + "丽": "lì", + "旗": "qí", + "編": "biān", + "编": "biān", + "胡": "hú", + "麗": "lì", + "穷": "qióng", + "窮": "qióng", + "雄": "xióng", + "玻": "bō", + "璃": "lí", + "剝": "bō", + "剥": "bō", + "粉": "fěn", + "艰": "jiān", + "艱": "jiān", + "零": "líng", + "肩": "jiān", + "云": "yún", + "挑": "tiāo", + "混": "hùn", + "顆": "kē", + "颗": "kē", + "善": "shàn", + "戏": "xì", + "戲": "xì", + "鑽": "zuān", + "钻": "zuān", + "借": "jiè", + "偷": "tōu", + "均": "jūn", + "昨": "zuó", + "舞": "wǔ", + "頓": "dùn", + "顿": "dùn", + "施": "shī", + "洲": "zhōu", + "篇": "piān", + "厚": "hòu", + "陆": "lù", + "陸": "lù", + "傅": "fù", + "招": "zhāo", + "范": "fàn", + "醒": "xǐng", + "剩": "shèng", + "福": "fú", + "默": "mò", + "良": "liáng", + "警": "jǐng", + "躺": "tǎng", + "休": "xiū", + "升": "shēng", + "圆": "yuán", + "圓": "yuán", + "夏": "xià", + "夺": "duó", + "奪": "duó", + "恶": "è", + "惡": "è", + "纖": "xiān", + "纤": "xiān", + "俩": "liǎ", + "倆": "liǎ", + "亿": "yì", + "億": "yì", + "擦": "cā", + "盘": "pán", + "盤": "pán", + "茶": "chá", + "伯": "bó", + "免": "miǎn", + "弱": "ruò", + "征": "zhēng", + "遭": "zāo", + "控": "kòng", + "迅": "xùn", + "堂": "táng", + "岛": "dǎo", + "島": "dǎo", + "虎": "hǔ", + "鳥": "niǎo", + "鸟": "niǎo", + "鼻": "bí", + "齊": "qí", + "齐": "qí", + "忍": "rěn", + "灰": "huī", + "爆": "bào", + "威": "wēi", + "帽": "mào", + "毒": "dú", + "牲": "shēng", + "冒": "mào", + "牙": "yá", + "丝": "sī", + "液": "yè", + "絲": "sī", + "宽": "kuān", + "寬": "kuān", + "灵": "líng", + "靈": "líng", + "居": "jū", + "松": "sōng", + "訓": "xùn", + "训": "xùn", + "罪": "zuì", + "炮": "pào", + "粗": "cū", + "罵": "mà", + "膀": "bǎng", + "若": "ruò", + "骂": "mà", + "圈": "quān", + "孔": "kǒng", + "貴": "guì", + "贵": "guì", + "扬": "yáng", + "揚": "yáng", + "楼": "lóu", + "樓": "lóu", + "献": "xiàn", + "獻": "xiàn", + "縮": "suō", + "缩": "suō", + "份": "fèn", + "紡": "fǎng", + "纺": "fǎng", + "胸": "xiōng", + "輛": "liàng", + "辆": "liàng", + "途": "tú", + "炉": "lú", + "爐": "lú", + "渡": "dù", + "耳": "ěr", + "倾": "qīng", + "傾": "qīng", + "涂": "tú", + "票": "piào", + "菌": "jūn", + "壮": "zhuàng", + "壯": "zhuàng", + "播": "bō", + "械": "xiè", + "拖": "tuō", + "职": "zhí", + "職": "zhí", + "克": "kè", + "帐": "zhàng", + "帳": "zhàng", + "挤": "jǐ", + "擠": "jǐ", + "秋": "qiū", + "括": "kuò", + "索": "suǒ", + "肚": "dù", + "插": "chā", + "棵": "kē", + "湿": "shī", + "濕": "shī", + "謂": "wèi", + "谓": "wèi", + "麻": "má", + "尾": "wěi", + "阿": "ā", + "尖": "jiān", + "慌": "huāng", + "梁": "liáng", + "涌": "yǒng", + "盆": "pén", + "蛋": "dàn", + "趣": "qù", + "冰": "bīng", + "怒": "nù", + "咬": "yǎo", + "財": "cái", + "财": "cái", + "避": "bì", + "累": "lèi", + "辩": "biàn", + "辯": "biàn", + "曲": "qū", + "磨": "mó", + "逃": "táo", + "餓": "è", + "饿": "è", + "承": "chéng", + "疑": "yí", + "刺": "cì", + "探": "tàn", + "糊": "hú", + "肥": "féi", + "贊": "zàn", + "赞": "zàn", + "弯": "wān", + "彎": "wān", + "徒": "tú", + "香": "xiāng", + "付": "fù", + "腰": "yāo", + "愤": "fèn", + "憤": "fèn", + "扩": "kuò", + "擴": "kuò", + "暖": "nuǎn", + "吨": "dūn", + "噸": "dūn", + "阻": "zǔ", + "介": "jiè", + "柴": "chái", + "獲": "huò", + "紹": "shào", + "绍": "shào", + "获": "huò", + "藏": "cáng", + "緩": "huǎn", + "缓": "huǎn", + "隔": "gé", + "奔": "bēn", + "秘": "mì", + "偏": "piān", + "叹": "tàn", + "嘆": "tàn", + "窝": "wō", + "窩": "wō", + "净": "jìng", + "晨": "chén", + "淨": "jìng", + "稳": "wěn", + "穩": "wěn", + "詩": "shī", + "诗": "shī", + "喂": "wèi", + "暴": "bào", + "殖": "zhí", + "潮": "cháo", + "协": "xié", + "協": "xié", + "登": "dēng", + "迷": "mí", + "壁": "bì", + "毕": "bì", + "畢": "bì", + "浮": "fú", + "紛": "fēn", + "纷": "fēn", + "闊": "kuò", + "阔": "kuò", + "阴": "yīn", + "附": "fù", + "陰": "yīn", + "井": "jǐng", + "哼": "hēng", + "巧": "qiǎo", + "拼": "pīn", + "榮": "róng", + "滚": "gǔn", + "滾": "gǔn", + "荣": "róng", + "厉": "lì", + "厲": "lì", + "异": "yì", + "異": "yì", + "麥": "mài", + "麦": "mài", + "寒": "hán", + "惯": "guàn", + "慣": "guàn", + "谷": "gǔ", + "丟": "diū", + "丢": "diū", + "培": "péi", + "宇": "yǔ", + "泛": "fàn", + "肃": "sù", + "肅": "sù", + "載": "zài", + "载": "zài", + "录": "lù", + "舒": "shū", + "錄": "lù", + "健": "jiàn", + "婆": "pó", + "搬": "bān", + "禁": "jìn", + "寻": "xún", + "尋": "xún", + "灌": "guàn", + "补": "bǔ", + "補": "bǔ", + "駝": "tuó", + "驼": "tuó", + "促": "cù", + "刷": "shuā", + "扑": "pū", + "撲": "pū", + "析": "xī", + "珠": "zhū", + "愈": "yù", + "旅": "lǚ", + "跃": "yuè", + "躍": "yuè", + "凝": "níng", + "彩": "cǎi", + "拔": "bá", + "袖": "xiù", + "幕": "mù", + "庭": "tíng", + "戴": "dài", + "援": "yuán", + "航": "háng", + "呆": "dāi", + "挖": "wā", + "杆": "gān", + "沟": "gōu", + "溝": "gōu", + "猿": "yuán", + "瓜": "guā", + "凡": "fán", + "吓": "xià", + "嚇": "xià", + "迎": "yíng", + "凭": "píng", + "憑": "píng", + "扫": "sǎo", + "掃": "sǎo", + "騎": "qí", + "骑": "qí", + "冻": "dòng", + "凍": "dòng", + "扎": "zhā", + "操": "cāo", + "箱": "xiāng", + "純": "chún", + "纯": "chún", + "聞": "wén", + "闻": "wén", + "仔": "zǐ", + "績": "jī", + "绩": "jì", + "訊": "xùn", + "讯": "xùn", + "踏": "tà", + "顏": "yán", + "颜": "yán", + "序": "xù", + "恨": "hèn", + "抢": "qiǎng", + "搶": "qiǎng", + "横": "héng", + "橫": "héng", + "疯": "fēng", + "瘋": "fēng", + "眉": "méi", + "宙": "zhòu", + "凉": "liáng", + "卷": "juǎn", + "夢": "mèng", + "梦": "mèng", + "氧": "yǎng", + "涼": "liáng", + "繁": "fán", + "距": "jù", + "銅": "tóng", + "铜": "tóng", + "仗": "zhàng", + "割": "gē", + "损": "sǔn", + "損": "sǔn", + "摄": "shè", + "摔": "shuāi", + "攝": "shè", + "瓶": "píng", + "悲": "bēi", + "昏": "hūn", + "疼": "téng", + "繩": "shéng", + "绳": "shéng", + "豆": "dòu", + "烂": "làn", + "烦": "fán", + "煩": "fán", + "爛": "làn", + "蓝": "lán", + "藍": "lán", + "訂": "dìng", + "订": "dìng", + "侧": "cè", + "側": "cè", + "巩": "gǒng", + "慮": "lǜ", + "虑": "lǜ", + "軟": "ruǎn", + "软": "ruǎn", + "鞏": "gǒng", + "匆": "cōng", + "域": "yù", + "尺": "chǐ", + "貼": "tiē", + "賽": "sài", + "贴": "tiē", + "赛": "sài", + "躲": "duǒ", + "剧": "jù", + "劇": "jù", + "役": "yì", + "恰": "qià", + "惟": "wéi", + "狠": "hěn", + "薄": "báo", + "释": "shì", + "釋": "shì", + "駛": "shǐ", + "驶": "shǐ", + "俺": "ǎn", + "兄": "xiōng", + "尊": "zūn", + "幅": "fú", + "拥": "yōng", + "授": "shòu", + "擁": "yōng", + "杯": "bēi", + "謀": "móu", + "谋": "móu", + "劝": "quàn", + "勸": "quàn", + "博": "bó", + "仪": "yí", + "儀": "yí", + "捧": "pěng", + "睁": "zhēng", + "睜": "zhēng", + "網": "wǎng", + "网": "wǎng", + "触": "chù", + "觸": "chù", + "腾": "téng", + "騰": "téng", + "匪": "fěi", + "夹": "jiā", + "夾": "jiā", + "抖": "dǒu", + "揭": "jiē", + "稍": "shāo", + "稼": "jià", + "腐": "fǔ", + "閉": "bì", + "闭": "bì", + "浓": "nóng", + "濃": "nóng", + "胞": "bāo", + "脈": "mài", + "脉": "mài", + "駱": "luò", + "骆": "luò", + "刑": "xíng", + "惜": "xī", + "皱": "zhòu", + "皺": "zhòu", + "监": "jiān", + "監": "jiān", + "脏": "zàng", + "臟": "zàng", + "蒸": "zhēng", + "貧": "pín", + "贫": "pín", + "鍋": "guō", + "锅": "guō", + "波": "bō", + "炸": "zhà", + "礼": "lǐ", + "禮": "lǐ", + "私": "sī", + "繞": "rào", + "绕": "rào", + "塑": "sù", + "磁": "cí", + "违": "wéi", + "違": "wéi", + "丈": "zhàng", + "玉": "yù", + "茫": "máng", + "吐": "tǔ", + "喷": "pēn", + "噴": "pēn", + "废": "fèi", + "廢": "fèi", + "怜": "lián", + "恢": "huī", + "悉": "xī", + "憐": "lián", + "挨": "āi", + "敲": "qiāo", + "淡": "dàn", + "尤": "yóu", + "忆": "yì", + "憶": "yì", + "災": "zāi", + "灾": "zāi", + "蜜": "mì", + "啥": "shá", + "恐": "kǒng", + "述": "shù", + "隐": "yǐn", + "隱": "yǐn", + "残": "cán", + "殘": "cán", + "額": "é", + "额": "é", + "亩": "mǔ", + "旋": "xuán", + "污": "wū", + "甲": "jiǎ", + "畝": "mǔ", + "胆": "dǎn", + "膽": "dǎn", + "蹲": "dūn", + "迟": "chí", + "遲": "chí", + "乘": "chéng", + "伴": "bàn", + "掏": "tāo", + "縫": "fèng", + "缝": "fèng", + "刮": "guā", + "椅": "yǐ", + "串": "chuàn", + "埋": "mái", + "抵": "dǐ", + "捉": "zhuō", + "秒": "miǎo", + "乏": "fá", + "喔": "ō", + "噢": "ō", + "坡": "pō", + "捕": "bǔ", + "添": "tiān", + "牺": "xī", + "犧": "xī", + "粒": "lì", + "舍": "shě", + "允": "yǔn", + "哇": "wa", + "柜": "guì", + "酸": "suān", + "寄": "jì", + "扔": "rēng", + "托": "tuō", + "措": "cuò", + "狂": "kuáng", + "遗": "yí", + "遺": "yí", + "伏": "fú", + "兔": "tù", + "勤": "qín", + "珍": "zhēn", + "糟": "zāo", + "輝": "huī", + "辉": "huī", + "拾": "shí", + "殊": "shū", + "浑": "hún", + "渾": "hún", + "滴": "dī", + "典": "diǎn", + "漠": "mò", + "猜": "cāi", + "障": "zhàng", + "唇": "chún", + "壳": "ké", + "峡": "xiá", + "峽": "xiá", + "德": "dé", + "忿": "fèn", + "撞": "zhuàng", + "棒": "bàng", + "殼": "ké", + "滑": "huá", + "牵": "qiān", + "牽": "qiān", + "盛": "shèng", + "糖": "táng", + "貢": "gòng", + "贡": "gòng", + "哟": "yō", + "喲": "yō", + "宜": "yí", + "敬": "jìng", + "斜": "xié", + "暂": "zàn", + "暫": "zàn", + "歼": "jiān", + "殲": "jiān", + "竹": "zhú", + "笼": "lóng", + "籠": "lóng", + "聪": "cōng", + "聰": "cōng", + "蜂": "fēng", + "騙": "piàn", + "骗": "piàn", + "扭": "niǔ", + "詳": "xiáng", + "详": "xiáng", + "貌": "mào", + "辟": "pì", + "亡": "wáng", + "峰": "fēng", + "励": "lì", + "勵": "lì", + "归": "guī", + "歸": "guī", + "焊": "hàn", + "秀": "xiù", + "唤": "huàn", + "喚": "huàn", + "寸": "cùn", + "毀": "huǐ", + "毁": "huǐ", + "稻": "dào", + "緒": "xù", + "绪": "xù", + "脆": "cuì", + "銷": "xiāo", + "销": "xiāo", + "库": "kù", + "庫": "kù", + "渠": "qú", + "爹": "diē", + "祝": "zhù", + "貫": "guàn", + "贯": "guàn", + "雷": "léi", + "坑": "kēng", + "蒙": "méng", + "辛": "xīn", + "遵": "zūn", + "飄": "piāo", + "飘": "piāo", + "婚": "hūn", + "披": "pī", + "胃": "wèi", + "趟": "tàng", + "逼": "bī", + "閑": "xián", + "闲": "xián", + "嚷": "rǎng", + "垂": "chuí", + "塞": "sāi", + "娃": "wá", + "扯": "chě", + "狼": "láng", + "鍛": "duàn", + "锻": "duàn", + "凳": "dèng", + "卵": "luǎn", + "炕": "kàng", + "箭": "jiàn", + "肤": "fū", + "膚": "fū", + "跡": "jī", + "輩": "bèi", + "辈": "bèi", + "迹": "jì", + "匠": "jiàng", + "巾": "jīn", + "洁": "jié", + "涨": "zhǎng", + "漲": "zhǎng", + "潔": "jié", + "猴": "hóu", + "耗": "hào", + "臂": "bì", + "虚": "xū", + "虛": "xū", + "陷": "xiàn", + "吵": "chǎo", + "咳": "hāi", + "搭": "dā", + "森": "sēn", + "漂": "piào", + "狱": "yù", + "獄": "yù", + "疗": "liáo", + "療": "liáo", + "皇": "huáng", + "翅": "chì", + "脾": "pí", + "鈴": "líng", + "铃": "líng", + "雾": "wù", + "霧": "wù", + "飽": "bǎo", + "饱": "bǎo", + "尚": "shàng", + "拣": "jiǎn", + "振": "zhèn", + "掩": "yǎn", + "揀": "jiǎn", + "歇": "xiē", + "牧": "mù", + "番": "fān", + "符": "fú", + "趁": "chèn", + "挡": "dǎng", + "擋": "dǎng", + "晓": "xiǎo", + "曉": "xiǎo", + "猪": "zhū", + "綱": "gāng", + "纲": "gāng", + "舅": "jiù", + "豬": "zhū", + "迈": "mài", + "递": "dì", + "遞": "dì", + "邁": "mài", + "壤": "rǎng", + "撤": "chè", + "浅": "qiǎn", + "淺": "qiǎn", + "瘦": "shòu", + "肠": "cháng", + "腸": "cháng", + "塘": "táng", + "塵": "chén", + "妙": "miào", + "尘": "chén", + "砍": "kǎn", + "碑": "bēi", + "焦": "jiāo", + "衡": "héng", + "齒": "chǐ", + "齿": "chǐ", + "剂": "jì", + "剑": "jiàn", + "劍": "jiàn", + "劑": "jì", + "匹": "pǐ", + "摘": "zhāi", + "竞": "jìng", + "競": "jìng", + "凶": "xiōng", + "售": "shòu", + "堵": "dǔ", + "康": "kāng", + "拱": "gǒng", + "漆": "qī", + "疲": "pí", + "盒": "hé", + "紗": "shā", + "纱": "shā", + "坦": "tǎn", + "斯": "sī", + "杨": "yáng", + "楊": "yáng", + "泡": "pào", + "盟": "méng", + "瞪": "dèng", + "緣": "yuán", + "缘": "yuán", + "苹": "píng", + "蘋": "píng", + "轟": "hōng", + "轰": "hōng", + "逗": "dòu", + "享": "xiǎng", + "喘": "chuǎn", + "嘿": "hēi", + "挣": "zhēng", + "掙": "zhēng", + "棚": "péng", + "签": "qiān", + "簽": "qiān", + "龍": "lóng", + "龙": "lóng", + "宿": "sù", + "悶": "mèn", + "泼": "pō", + "溉": "gài", + "潑": "pō", + "甜": "tián", + "舱": "cāng", + "艙": "cāng", + "闷": "mèn", + "餅": "bǐng", + "饼": "bǐng", + "愉": "yú", + "捏": "niē", + "棍": "gùn", + "盼": "pàn", + "篮": "lán", + "籃": "lán", + "芦": "lú", + "蘆": "lú", + "鉛": "qiān", + "铅": "qiān", + "匯": "huì", + "奴": "nú", + "宫": "gōng", + "宮": "gōng", + "汇": "huì", + "炭": "tàn", + "版": "bǎn", + "牌": "pái", + "窑": "yáo", + "窯": "yáo", + "聚": "jù", + "脖": "bó", + "訪": "fǎng", + "访": "fǎng", + "隶": "lì", + "隸": "lì", + "咐": "fù", + "摊": "tān", + "攤": "tān", + "昆": "kūn", + "桶": "tǒng", + "池": "chí", + "猎": "liè", + "獵": "liè", + "碍": "ài", + "礙": "ài", + "臭": "chòu", + "詞": "cí", + "词": "cí", + "軌": "guǐ", + "轨": "guǐ", + "釣": "diào", + "钓": "diào", + "顫": "chàn", + "颤": "chàn", + "亏": "kuī", + "仇": "chóu", + "择": "zé", + "擇": "zé", + "智": "zhì", + "苗": "miáo", + "虧": "kuī", + "鋒": "fēng", + "锋": "fēng", + "仰": "yǎng", + "屆": "jiè", + "届": "jiè", + "岗": "gǎng", + "岩": "yán", + "岭": "lǐng", + "崗": "gǎng", + "嶺": "lǐng", + "慰": "wèi", + "抄": "chāo", + "盐": "yán", + "譯": "yì", + "译": "yì", + "鹽": "yán", + "丛": "cóng", + "乌": "wū", + "凑": "còu", + "厘": "lí", + "叢": "cóng", + "奖": "jiǎng", + "妻": "qī", + "径": "jìng", + "徑": "jìng", + "悟": "wù", + "欠": "qiàn", + "湊": "còu", + "烏": "wū", + "獎": "jiǎng", + "苍": "cāng", + "荷": "hé", + "蒼": "cāng", + "輯": "jí", + "辑": "jí", + "陪": "péi", + "叛": "pàn", + "捞": "lāo", + "撈": "lāo", + "撒": "sā", + "柱": "zhù", + "株": "zhū", + "核": "hé", + "润": "rùn", + "漏": "lòu", + "潤": "rùn", + "瓷": "cí", + "糾": "jiū", + "纠": "jiū", + "蛇": "shé", + "鎖": "suǒ", + "锁": "suǒ", + "估": "gū", + "傲": "ào", + "厌": "yàn", + "厭": "yàn", + "宗": "zōng", + "扶": "fú", + "捆": "kǔn", + "荡": "dàng", + "蕩": "dàng", + "蚀": "shí", + "蝕": "shí", + "裂": "liè", + "驕": "jiāo", + "骄": "jiāo", + "幼": "yòu", + "拨": "bō", + "挽": "wǎn", + "掀": "xiān", + "撥": "bō", + "銳": "ruì", + "锐": "ruì", + "鳴": "míng", + "鸣": "míng", + "款": "kuǎn", + "盯": "dīng", + "胳": "gē", + "偶": "ǒu", + "寂": "jì", + "屈": "qū", + "恳": "kěn", + "懇": "kěn", + "晃": "huǎng", + "歪": "wāi", + "眯": "mī", + "瞇": "mī", + "秧": "yāng", + "稿": "gǎo", + "綜": "zōng", + "综": "zōng", + "踩": "cǎi", + "鯨": "jīng", + "鲸": "jīng", + "吼": "hǒu", + "嗓": "sǎng", + "扁": "biǎn", + "朴": "pǔ", + "欣": "xīn", + "莫": "mò", + "傻": "shǎ", + "幻": "huàn", + "扣": "kòu", + "拢": "lǒng", + "掠": "lüè", + "攏": "lǒng", + "榴": "liú", + "溶": "róng", + "滩": "tān", + "灘": "tān", + "牢": "láo", + "猫": "māo", + "腔": "qiāng", + "蚕": "cán", + "蝗": "huáng", + "蠶": "cán", + "裤": "kù", + "褲": "kù", + "貓": "māo", + "跨": "kuà", + "霜": "shuāng", + "冶": "yě", + "咽": "yàn", + "宅": "zhái", + "搜": "sōu", + "晴": "qíng", + "遮": "zhē", + "启": "qǐ", + "啟": "qǐ", + "彼": "bǐ", + "抹": "mǒ", + "搁": "gē", + "擱": "gē", + "敏": "mǐn", + "漫": "màn", + "码": "mǎ", + "碼": "mǎ", + "筋": "jīn", + "鍵": "jiàn", + "键": "jiàn", + "厅": "tīng", + "吊": "diào", + "廳": "tīng", + "拒": "jù", + "旱": "hàn", + "桃": "táo", + "欺": "qī", + "燕": "yàn", + "琴": "qín", + "舌": "shé", + "蔽": "bì", + "袄": "ǎo", + "襖": "ǎo", + "釘": "dīng", + "钉": "dīng", + "駕": "jià", + "驾": "jià", + "丘": "qiū", + "审": "shěn", + "審": "shěn", + "币": "bì", + "幣": "bì", + "愣": "lèng", + "拦": "lán", + "摧": "cuī", + "撕": "sī", + "攔": "lán", + "浇": "jiāo", + "澆": "jiāo", + "賞": "shǎng", + "赏": "shǎng", + "鴉": "yā", + "鸦": "yā", + "伞": "sǎn", + "傘": "sǎn", + "晶": "jīng", + "涉": "shè", + "犹": "yóu", + "猶": "yóu", + "蛙": "wā", + "丫": "yā", + "僚": "liáo", + "哏": "gén", + "嗡": "wēng", + "宪": "xiàn", + "憲": "xiàn", + "描": "miáo", + "朗": "lǎng", + "柔": "róu", + "橘": "jú", + "瞎": "xiā", + "稀": "xī", + "肝": "gān", + "裳": "shang", + "隆": "lóng", + "頑": "wán", + "顽": "wán", + "驢": "lǘ", + "驴": "lǘ", + "倡": "chàng", + "哀": "āi", + "堤": "dī", + "姨": "yí", + "崇": "chóng", + "庙": "miào", + "廟": "miào", + "延": "yán", + "汤": "tāng", + "湯": "tāng", + "碳": "tàn", + "童": "tóng", + "耕": "gēng", + "跪": "guì", + "辫": "biàn", + "辮": "biàn", + "闖": "chuǎng", + "闯": "chuǎng", + "頗": "pō", + "颇": "pō", + "勃": "bó", + "哗": "huā", + "嘩": "huā", + "嫁": "jià", + "孤": "gū", + "拳": "quán", + "晒": "shài", + "栽": "zāi", + "洒": "sǎ", + "耀": "yào", + "胀": "zhàng", + "胁": "xié", + "脅": "xié", + "脹": "zhàng", + "膜": "mó", + "荒": "huāng", + "亭": "tíng", + "咧": "liě", + "填": "tián", + "妥": "tuǒ", + "帘": "lián", + "患": "huàn", + "截": "jié", + "抑": "yì", + "攀": "pān", + "梅": "méi", + "烛": "zhú", + "燭": "zhú", + "督": "dū", + "逢": "féng", + "魔": "mó", + "伐": "fá", + "媳": "xí", + "悬": "xuán", + "懸": "xuán", + "戚": "qī", + "煮": "zhǔ", + "盗": "dào", + "盜": "dào", + "綁": "bǎng", + "绑": "bǎng", + "肺": "fèi", + "侦": "zhēn", + "俗": "sú", + "偵": "zhēn", + "哨": "shào", + "喉": "hóu", + "岂": "qǐ", + "帜": "zhì", + "幟": "zhì", + "庆": "qìng", + "弃": "qì", + "惨": "cǎn", + "慘": "cǎn", + "慶": "qìng", + "抛": "pāo", + "拋": "pāo", + "末": "mò", + "棄": "qì", + "浸": "jìn", + "港": "gǎng", + "眨": "zhǎ", + "租": "zū", + "窜": "cuàn", + "竄": "cuàn", + "誠": "chéng", + "诚": "chéng", + "豈": "qǐ", + "醉": "zuì", + "刊": "kān", + "墨": "mò", + "桩": "zhuāng", + "樁": "zhuāng", + "炎": "yán", + "盏": "zhǎn", + "盞": "zhǎn", + "肖": "xiào", + "踢": "tī", + "錦": "jǐn", + "锦": "jǐn", + "啪": "pā", + "塔": "tǎ", + "惹": "rě", + "柳": "liǔ", + "筐": "kuāng", + "紫": "zǐ", + "罩": "zhào", + "萄": "táo", + "葡": "pú", + "貝": "bèi", + "贝": "bèi", + "辨": "biàn", + "顛": "diān", + "颠": "diān", + "伪": "wěi", + "偽": "wěi", + "冤": "yuān", + "厨": "chú", + "吩": "fēn", + "妄": "wàng", + "姿": "zī", + "屁": "pì", + "廚": "chú", + "愁": "chóu", + "晌": "shǎng", + "渴": "kě", + "溜": "liū", + "甩": "shuǎi", + "眠": "mián", + "粥": "zhōu", + "綸": "lún", + "纶": "lún", + "蝉": "chán", + "蟬": "chán", + "覽": "lǎn", + "览": "lǎn", + "返": "fǎn", + "閥": "fá", + "阀": "fá", + "鞭": "biān", + "頻": "pín", + "频": "pín", + "仓": "cāng", + "倉": "cāng", + "傍": "bàng", + "壶": "hú", + "壺": "hú", + "怨": "yuàn", + "汗": "hàn", + "泉": "quán", + "窄": "zhǎi", + "紋": "wén", + "纹": "wén", + "跌": "diē", + "喽": "lóu", + "嘍": "lóu", + "坟": "fén", + "墳": "fén", + "扛": "káng", + "扮": "bàn", + "洪": "hóng", + "瓦": "wǎ", + "秩": "zhì", + "脂": "zhī", + "虾": "xiā", + "蝦": "xiā", + "衬": "chèn", + "袭": "xí", + "裹": "guǒ", + "襯": "chèn", + "襲": "xí", + "諒": "liàng", + "谅": "liàng", + "魂": "hún", + "乙": "yǐ", + "倘": "tǎng", + "卧": "wò", + "矮": "ǎi", + "筒": "tǒng", + "膊": "bó", + "臥": "wò", + "揉": "róu", + "昂": "áng", + "栏": "lán", + "欄": "lán", + "疾": "jí", + "痕": "hén", + "砖": "zhuān", + "磚": "zhuān", + "膨": "péng", + "餐": "cān", + "兜": "dōu", + "夸": "kuā", + "崖": "yá", + "拆": "chāi", + "斧": "fǔ", + "欲": "yù", + "沫": "mò", + "涡": "wō", + "渦": "wō", + "縱": "zòng", + "纵": "zòng", + "肌": "jī", + "胖": "pàng", + "趴": "pā", + "飲": "yǐn", + "饮": "yǐn", + "齡": "líng", + "龄": "líng", + "丹": "dān", + "勾": "gōu", + "嘻": "xī", + "御": "yù", + "戒": "jiè", + "拴": "shuān", + "撐": "chēng", + "撑": "chēng", + "朽": "xiǔ", + "甘": "gān", + "袜": "wà", + "袱": "fú", + "裁": "cái", + "襪": "wà", + "譬": "pì", + "鉤": "gōu", + "鋁": "lǚ", + "钩": "gōu", + "铝": "lǚ", + "鼠": "shǔ", + "催": "cuī", + "咦": "yí", + "拧": "níng", + "搅": "jiǎo", + "擰": "níng", + "攪": "jiǎo", + "淹": "yān", + "渔": "yú", + "漁": "yú", + "熊": "xióng", + "盲": "máng", + "筷": "kuài", + "緯": "wěi", + "纬": "wěi", + "購": "gòu", + "购": "gòu", + "鴨": "yā", + "鸭": "yā", + "予": "yǔ", + "兼": "jiān", + "兽": "shòu", + "呈": "chéng", + "哄": "hōng", + "娶": "qǔ", + "恆": "héng", + "恒": "héng", + "慧": "huì", + "梯": "tī", + "殿": "diàn", + "氏": "shì", + "淋": "lín", + "溪": "xī", + "獸": "shòu", + "罐": "guàn", + "蚁": "yǐ", + "蚂": "mǎ", + "蜡": "là", + "螞": "mǎ", + "蟻": "yǐ", + "蠟": "là", + "誕": "dàn", + "诞": "dàn", + "逮": "dǎi", + "飾": "shì", + "饰": "shì", + "剪": "jiǎn", + "叠": "dié", + "嗽": "sòu", + "悔": "huǐ", + "槽": "cáo", + "疊": "dié", + "碧": "bì", + "繪": "huì", + "绘": "huì", + "耸": "sǒng", + "聳": "sǒng", + "蝇": "yíng", + "蠅": "yíng", + "豫": "yù", + "蹬": "dēng", + "軸": "zhóu", + "轴": "zhóu", + "叮": "dīng", + "嘗": "cháng", + "圾": "jī", + "垃": "lā", + "垮": "kuǎ", + "尝": "cháng", + "慎": "shèn", + "沾": "zhān", + "潛": "qián", + "潜": "qián", + "皂": "zào", + "窃": "qiè", + "竊": "qiè", + "缸": "gāng", + "肢": "zhī", + "胎": "tāi", + "脊": "jǐ", + "膝": "xī", + "艳": "yàn", + "艷": "yàn", + "詫": "chà", + "诧": "chà", + "酷": "kù", + "雕": "diāo", + "霉": "méi", + "冈": "gāng", + "勉": "miǎn", + "吆": "yāo", + "嫌": "xián", + "岡": "gāng", + "巷": "xiàng", + "愧": "kuì", + "拌": "bàn", + "揪": "jiū", + "晰": "xī", + "泊": "pō", + "灿": "càn", + "燦": "càn", + "瓣": "bàn", + "症": "zhèng", + "胶": "jiāo", + "膠": "jiāo", + "豁": "huō", + "踱": "duó", + "閨": "guī", + "闺": "guī", + "隙": "xì", + "飢": "jī", + "饅": "mán", + "饥": "jī", + "馒": "mán", + "债": "zhài", + "債": "zhài", + "唰": "shuā", + "墩": "dūn", + "弓": "gōng", + "恥": "chǐ", + "旦": "dàn", + "李": "lǐ", + "烤": "kǎo", + "熄": "xī", + "砸": "zá", + "粪": "fèn", + "糞": "fèn", + "耻": "chǐ", + "誉": "yù", + "譽": "yù", + "貿": "mào", + "贸": "mào", + "酱": "jiàng", + "醬": "jiàng", + "鑄": "zhù", + "铸": "zhù", + "飼": "sì", + "饲": "sì", + "亦": "yì", + "仙": "xiān", + "哧": "chī", + "嘱": "zhǔ", + "囑": "zhǔ", + "妨": "fáng", + "婴": "yīng", + "嬰": "yīng", + "寞": "mò", + "押": "yā", + "斥": "chì", + "框": "kuàng", + "爽": "shuǎng", + "甭": "béng", + "畜": "chù", + "癌": "ái", + "硫": "liú", + "笨": "bèn", + "籍": "jí", + "芒": "máng", + "蝴": "hú", + "蝶": "dié", + "袍": "páo", + "豪": "háo", + "邻": "lín", + "鄰": "lín", + "頁": "yè", + "页": "yè", + "馳": "chí", + "驰": "chí", + "倚": "yǐ", + "僵": "jiāng", + "凿": "záo", + "勻": "yún", + "匀": "yún", + "君": "jūn", + "宴": "yàn", + "宵": "xiāo", + "崭": "zhǎn", + "嶄": "zhǎn", + "扇": "shàn", + "枕": "zhěn", + "枯": "kū", + "渗": "shèn", + "滲": "shèn", + "焰": "yàn", + "瞅": "chǒu", + "縛": "fù", + "缚": "fù", + "蛛": "zhū", + "蜘": "zhī", + "赤": "chì", + "迁": "qiān", + "遷": "qiān", + "鑿": "záo", + "埃": "āi", + "慨": "kǎi", + "挫": "cuò", + "淘": "táo", + "渣": "zhā", + "砂": "shā", + "耽": "dān", + "苏": "sū", + "蔬": "shū", + "蘇": "sū", + "訝": "yà", + "讶": "yà", + "躁": "zào", + "鉴": "jiàn", + "鑒": "jiàn", + "雀": "què", + "駐": "zhù", + "驻": "zhù", + "侮": "wǔ", + "吁": "xū", + "呸": "pēi", + "啾": "jiū", + "塌": "tā", + "循": "xún", + "怖": "bù", + "扰": "rǎo", + "擾": "rǎo", + "朦": "méng", + "朧": "lóng", + "燥": "zào", + "瞒": "mán", + "瞞": "mán", + "纏": "chán", + "缠": "chán", + "胧": "lóng", + "苇": "wěi", + "葦": "wěi", + "診": "zhěn", + "诊": "zhěn", + "辽": "liáo", + "遼": "liáo", + "邮": "yóu", + "郵": "yóu", + "陌": "mò", + "丑": "chǒu", + "俘": "fú", + "凸": "tū", + "凹": "āo", + "刹": "shā", + "剎": "shā", + "嗨": "hāi", + "宏": "hóng", + "懒": "lǎn", + "懶": "lǎn", + "扒": "bā", + "梢": "shāo", + "狭": "xiá", + "狹": "xiá", + "碱": "jiǎn", + "芽": "yá", + "虏": "lǔ", + "虜": "lǔ", + "賊": "zéi", + "贼": "zéi", + "輔": "fǔ", + "輻": "fú", + "辅": "fǔ", + "辐": "fú", + "逝": "shì", + "陡": "dǒu", + "陵": "líng", + "頌": "sòng", + "颂": "sòng", + "鹼": "jiǎn", + "黏": "nián", + "丙": "bǐng", + "吞": "tūn", + "哆": "duō", + "嗦": "suō", + "夕": "xī", + "屿": "yǔ", + "嶼": "yǔ", + "州": "zhōu", + "梳": "shū", + "汹": "xiōng", + "洶": "xiōng", + "浊": "zhuó", + "淀": "diàn", + "澱": "diàn", + "濁": "zhuó", + "烫": "tàng", + "燙": "tàng", + "爪": "zhǎo", + "竖": "shù", + "竿": "gān", + "繃": "běng", + "绷": "bēng", + "翘": "qiào", + "翹": "qiào", + "艘": "sōu", + "蚊": "wén", + "誒": "éi", + "誼": "yì", + "诶": "éi", + "谊": "yì", + "豎": "shù", + "貪": "tān", + "贪": "tān", + "遣": "qiǎn", + "邀": "yāo", + "黎": "lí", + "乒": "pīng", + "乓": "pāng", + "佩": "pèi", + "华": "huá", + "堪": "kān", + "昼": "zhòu", + "晝": "zhòu", + "柄": "bǐng", + "橡": "xiàng", + "毯": "tǎn", + "潭": "tán", + "烁": "shuò", + "爍": "shuò", + "矩": "jǔ", + "耍": "shuǎ", + "聊": "liáo", + "膛": "táng", + "茂": "mào", + "華": "huá", + "賤": "jiàn", + "贱": "jiàn", + "鏟": "chǎn", + "铲": "chǎn", + "倦": "juàn", + "卡": "kǎ", + "卸": "xiè", + "嘲": "cháo", + "囪": "cōng", + "囱": "cōng", + "圣": "shèng", + "垄": "lǒng", + "垒": "lěi", + "壕": "háo", + "壘": "lěi", + "壟": "lǒng", + "寡": "guǎ", + "巢": "cháo", + "氛": "fēn", + "沥": "lì", + "瀝": "lì", + "炯": "jiǒng", + "甸": "diàn", + "聖": "shèng", + "荔": "lì", + "葫": "hú", + "賴": "lài", + "赖": "lài", + "蹦": "bèng", + "鑼": "luó", + "锣": "luó", + "鵲": "què", + "鹊": "què", + "乃": "nǎi", + "忠": "zhōng", + "恼": "nǎo", + "惱": "nǎo", + "拜": "bài", + "搀": "chān", + "攙": "chān", + "旺": "wàng", + "昧": "mèi", + "滋": "zī", + "磷": "lín", + "竭": "jié", + "絡": "luò", + "絨": "róng", + "绒": "róng", + "络": "luò", + "署": "shǔ", + "脯": "pú", + "覆": "fù", + "訟": "sòng", + "讼": "sòng", + "轎": "jiào", + "轿": "jiào", + "霸": "bà", + "馱": "tuó", + "驮": "tuó", + "亚": "yà", + "亞": "yà", + "劣": "liè", + "咙": "lóng", + "喃": "nán", + "嚨": "lóng", + "奏": "zòu", + "奠": "diàn", + "嫩": "nèn", + "尔": "ěr", + "徐": "xú", + "捡": "jiǎn", + "搏": "bó", + "撿": "jiǎn", + "枉": "wǎng", + "煌": "huáng", + "爾": "ěr", + "猩": "xīng", + "盔": "kuī", + "窟": "kū", + "窿": "lóng", + "粘": "zhān", + "紐": "niǔ", + "纽": "niǔ", + "罚": "fá", + "罰": "fá", + "衫": "shān", + "謙": "qiān", + "謬": "miù", + "谦": "qiān", + "谬": "miù", + "郊": "jiāo", + "頃": "qǐng", + "顷": "qǐng", + "駁": "bó", + "驳": "bó", + "储": "chǔ", + "僱": "gù", + "儲": "chǔ", + "剿": "jiǎo", + "卜": "bo", + "叭": "bā", + "喇": "lǎ", + "奉": "fèng", + "弥": "mí", + "彌": "mí", + "揍": "zòu", + "搂": "lǒu", + "摟": "lǒu", + "旬": "xún", + "杰": "jié", + "毅": "yì", + "氓": "máng", + "沸": "fèi", + "涤": "dí", + "滌": "dí", + "灶": "zào", + "磺": "huáng", + "萝": "luó", + "蘿": "luó", + "融": "róng", + "誓": "shì", + "賺": "zhuàn", + "赚": "zhuàn", + "辣": "là", + "鈔": "chāo", + "錫": "xī", + "钞": "chāo", + "锡": "xī", + "闡": "chǎn", + "阐": "chǎn", + "雇": "gù", + "乳": "rǔ", + "伺": "cì", + "剖": "pōu", + "吟": "yín", + "址": "zhǐ", + "坝": "bà", + "坯": "pī", + "垫": "diàn", + "堡": "bǎo", + "墊": "diàn", + "壩": "bà", + "怦": "pēng", + "恍": "huǎng", + "惭": "cán", + "慕": "mù", + "慚": "cán", + "拐": "guǎi", + "砌": "qì", + "綢": "chóu", + "绸": "chóu", + "羡": "xiàn", + "羨": "xiàn", + "肿": "zhǒng", + "腫": "zhǒng", + "腹": "fù", + "蓬": "péng", + "詢": "xún", + "諸": "zhū", + "询": "xún", + "诸": "zhū", + "賠": "péi", + "赔": "péi", + "踌": "chóu", + "躇": "chú", + "躊": "chóu", + "逻": "luó", + "邏": "luó", + "陈": "chén", + "陳": "chén", + "鵝": "é", + "鹅": "é", + "匾": "biǎn", + "墓": "mù", + "忧": "yōu", + "悅": "yuè", + "悦": "yuè", + "惑": "huò", + "憂": "yōu", + "捣": "dǎo", + "搓": "cuō", + "搗": "dǎo", + "档": "dàng", + "檔": "dàng", + "歉": "qiàn", + "泌": "mì", + "溅": "jiàn", + "澡": "zǎo", + "濺": "jiàn", + "磕": "kē", + "稅": "shuì", + "税": "shuì", + "篷": "péng", + "翼": "yì", + "蟀": "shuài", + "蟋": "xī", + "辞": "cí", + "辭": "cí", + "遙": "yáo", + "遥": "yáo", + "陶": "táo", + "饒": "ráo", + "饶": "ráo", + "丧": "sàng", + "俯": "fǔ", + "呗": "bei", + "唄": "bei", + "喪": "sàng", + "奈": "nài", + "宁": "níng", + "寧": "níng", + "帆": "fān", + "廓": "kuò", + "拟": "nǐ", + "捂": "wǔ", + "擬": "nǐ", + "氨": "ān", + "汞": "gǒng", + "淌": "tǎng", + "炒": "chǎo", + "煎": "jiān", + "繡": "xiù", + "绣": "xiù", + "艇": "tǐng", + "躬": "gōng", + "辱": "rǔ", + "酬": "chóu", + "醋": "cù", + "鋤": "chú", + "鏽": "xiù", + "锄": "chú", + "锈": "xiù", + "閱": "yuè", + "阅": "yuè", + "隧": "suì", + "雌": "cí", + "鞠": "jū", + "骼": "gé", + "佳": "jiā", + "冊": "cè", + "册": "cè", + "啸": "xiào", + "嘯": "xiào", + "姻": "yīn", + "孵": "fū", + "憾": "hàn", + "扳": "bān", + "敷": "fū", + "棋": "qí", + "涛": "tāo", + "濤": "tāo", + "熔": "róng", + "熬": "áo", + "狮": "shī", + "獅": "shī", + "畔": "pàn", + "疏": "shū", + "納": "nà", + "纳": "nà", + "腈": "jīng", + "膏": "gāo", + "舰": "jiàn", + "艦": "jiàn", + "蜓": "tíng", + "蜻": "qīng", + "謊": "huǎng", + "谎": "huǎng", + "趋": "qū", + "趨": "qū", + "軋": "yà", + "轧": "yà", + "頒": "bān", + "颁": "bān", + "騾": "luó", + "骡": "luó", + "鵪": "ān", + "鶉": "chún", + "鹌": "ān", + "鹑": "chún", + "侍": "shì", + "壽": "shòu", + "奸": "jiān", + "寿": "shòu", + "慈": "cí", + "捷": "jié", + "枣": "zǎo", + "棗": "zǎo", + "榜": "bǎng", + "毡": "zhān", + "氈": "zhān", + "汪": "wāng", + "津": "jīn", + "滔": "tāo", + "狡": "jiǎo", + "猾": "huá", + "申": "shēn", + "畏": "wèi", + "祥": "xiáng", + "穗": "suì", + "簇": "cù", + "翁": "wēng", + "茸": "róng", + "蚩": "chī", + "蠕": "rú", + "衙": "yá", + "謹": "jǐn", + "谨": "jǐn", + "閘": "zhá", + "闸": "zhá", + "驟": "zhòu", + "骤": "zhòu", + "乖": "guāi", + "咆": "páo", + "哮": "xiào", + "孙": "sūn", + "孫": "sūn", + "寇": "kòu", + "崩": "bēng", + "庞": "páng", + "憋": "biē", + "捎": "shāo", + "敞": "chǎng", + "晕": "yūn", + "暈": "yūn", + "柏": "bǎi", + "瀑": "pù", + "烘": "hōng", + "熏": "xūn", + "舀": "yǎo", + "荐": "jiàn", + "蔼": "ǎi", + "藹": "ǎi", + "衰": "shuāi", + "謠": "yáo", + "譏": "jī", + "讥": "jī", + "谣": "yáo", + "跺": "duò", + "逛": "guàng", + "鷹": "yīng", + "鹰": "yīng", + "龐": "páng", + "俱": "jù", + "厢": "xiāng", + "叙": "xù", + "咂": "zā", + "屹": "yì", + "廂": "xiāng", + "怯": "qiè", + "拘": "jū", + "携": "xié", + "摩": "mó", + "攜": "xié", + "敘": "xù", + "暢": "chàng", + "梗": "gěng", + "沃": "wò", + "滥": "làn", + "濫": "làn", + "狐": "hú", + "狸": "lí", + "琢": "zuó", + "畅": "chàng", + "眶": "kuàng", + "簸": "bǒ", + "粜": "tiào", + "糕": "gāo", + "糶": "tiào", + "絹": "juàn", + "綿": "mián", + "縷": "lǚ", + "绢": "juàn", + "绵": "mián", + "缕": "lǚ", + "菩": "pú", + "萤": "yíng", + "萨": "sà", + "薩": "sà", + "螢": "yíng", + "鍍": "dù", + "镀": "dù", + "劈": "pī", + "厦": "shà", + "咋": "zǎ", + "啃": "kěn", + "屉": "tì", + "屜": "tì", + "嵌": "qiàn", + "廈": "shà", + "徊": "huái", + "徘": "pái", + "捍": "hàn", + "撼": "hàn", + "斃": "bì", + "杉": "shān", + "毙": "bì", + "泳": "yǒng", + "浆": "jiāng", + "湾": "wān", + "漾": "yàng", + "漿": "jiāng", + "灣": "wān", + "煞": "shā", + "疙": "gē", + "瘩": "da", + "碌": "lù", + "磅": "bàng", + "粹": "cuì", + "繳": "jiǎo", + "缰": "jiāng", + "缴": "jiǎo", + "舶": "bó", + "茅": "máo", + "薪": "xīn", + "裕": "yù", + "鉗": "qián", + "鑲": "xiāng", + "钳": "qián", + "镶": "xiāng", + "韁": "jiāng", + "丁": "dīng", + "偎": "wēi", + "凄": "qī", + "凤": "fèng", + "凰": "huáng", + "叼": "diāo", + "姆": "mǔ", + "尿": "niào", + "弦": "xián", + "惕": "tì", + "惧": "jù", + "懼": "jù", + "挎": "kuà", + "撅": "juē", + "杜": "dù", + "桨": "jiǎng", + "槳": "jiǎng", + "樟": "zhāng", + "欧": "ōu", + "歐": "ōu", + "淒": "qī", + "淳": "chún", + "渺": "miǎo", + "珊": "shān", + "瑚": "hú", + "痒": "yǎng", + "瞥": "piē", + "砰": "pēng", + "硝": "xiāo", + "祸": "huò", + "禍": "huò", + "稚": "zhì", + "糙": "cāo", + "紳": "shēn", + "绅": "shēn", + "羽": "yǔ", + "舔": "tiǎn", + "葵": "kuí", + "蚓": "yǐn", + "蚯": "qiū", + "豺": "chái", + "郑": "zhèng", + "鄭": "zhèng", + "鐺": "dāng", + "铛": "dāng", + "鳳": "fèng", + "鵑": "juān", + "鹃": "juān", + "伶": "líng", + "咀": "jǔ", + "噪": "zào", + "嚼": "jué", + "娛": "yú", + "娱": "yú", + "屠": "tú", + "怔": "zhēng", + "惩": "chéng", + "懲": "chéng", + "捐": "juān", + "捶": "chuí", + "撩": "liāo", + "枚": "méi", + "枢": "shū", + "槛": "kǎn", + "樞": "shū", + "檻": "kǎn", + "漩": "xuán", + "碟": "dié", + "秤": "chèng", + "竽": "yú", + "笆": "bā", + "篱": "lí", + "籬": "lí", + "臣": "chén", + "茎": "jīng", + "莖": "jīng", + "蚜": "yá", + "蝙": "biān", + "蝠": "fú", + "褂": "guà", + "詭": "guǐ", + "诡": "guǐ", + "豹": "bào", + "賀": "hè", + "贺": "hè", + "蹄": "tí", + "鈾": "yóu", + "錘": "chuí", + "鍬": "qiāo", + "铀": "yóu", + "锤": "chuí", + "锹": "qiāo", + "雹": "báo", + "霞": "xiá", + "呃": "è", + "咕": "gū", + "哑": "yǎ", + "哺": "bǔ", + "唬": "hǔ", + "唾": "tuò", + "啞": "yǎ", + "嗅": "xiù", + "嗐": "hài", + "嘀": "dí", + "嘶": "sī", + "尸": "shī", + "屎": "shǐ", + "悠": "yōu", + "惋": "wǎn", + "愕": "è", + "抿": "mǐn", + "掷": "zhì", + "搔": "sāo", + "搪": "táng", + "擲": "zhì", + "旷": "kuàng", + "曠": "kuàng", + "桅": "wéi", + "梨": "lí", + "榕": "róng", + "泣": "qì", + "泵": "bèng", + "涕": "tì", + "猬": "wèi", + "瑩": "yíng", + "疆": "jiāng", + "矗": "chù", + "硅": "guī", + "綴": "zhuì", + "繭": "jiǎn", + "缀": "zhuì", + "腮": "sāi", + "芯": "xīn", + "茧": "jiǎn", + "莹": "yíng", + "蕉": "jiāo", + "蕴": "yùn", + "藤": "téng", + "蘊": "yùn", + "蝟": "wèi", + "誣": "wū", + "誦": "sòng", + "諷": "fěng", + "讽": "fěng", + "诬": "wū", + "诵": "sòng", + "販": "fàn", + "贩": "fàn", + "蹭": "cèng", + "鈷": "gǔ", + "鋸": "jù", + "鎂": "měi", + "鐮": "lián", + "钴": "gǔ", + "锯": "jù", + "镁": "měi", + "镰": "lián", + "隘": "ài", + "髦": "máo", + "鱷": "è", + "鳄": "è", + "鸝": "lí", + "鹂": "lí", + "丸": "wán", + "伊": "yī", + "勒": "lēi", + "勘": "kān", + "匙": "shi", + "卑": "bēi", + "吉": "jí" + } +} diff --git a/scripts/generate-pinyin.mjs b/scripts/generate-pinyin.mjs new file mode 100644 index 00000000..a7d32e2e --- /dev/null +++ b/scripts/generate-pinyin.mjs @@ -0,0 +1,204 @@ +#!/usr/bin/env node +/** + * Derive `assets/pinyin.json` from the Unicode Unihan database. + * + * `pypinyin` was the writer's only optional Python dependency. The Node core + * ships a small, reviewable table instead, built from a single upstream source + * with a documented selection rule: + * + * characters = the 3500 most frequent Han characters according to the summed + * `kHanyuPinlu` corpus frequency in Unihan_Readings.txt; the reading is + * `kMandarin` (first reading) or, failing that, the highest-frequency + * `kHanyuPinlu` reading. + * + * Usage: + * node scripts/generate-pinyin.mjs # download Unihan.zip, extract, write the asset + * node scripts/generate-pinyin.mjs --source # use a local Unihan_Readings.txt + * node scripts/generate-pinyin.mjs --check # fail when the committed asset is stale + * node scripts/generate-pinyin.mjs --limit 0 # keep every covered character + * + * The output is deterministic (code-point order, no wall-clock timestamps): the + * `source_date` comes from the Unihan file header, so `--check` is a real drift + * gate and can run in CI. + */ + +import { spawnSync } from "node:child_process"; +import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; + +const here = dirname(fileURLToPath(import.meta.url)); +const repoRoot = resolve(here, ".."); +const DEFAULT_OUT = join(repoRoot, "assets", "pinyin.json"); +const DEFAULT_LIMIT = 3500; +const UNIHAN_URL = "https://www.unicode.org/Public/UCD/latest/ucd/Unihan.zip"; +const UNICODE_LICENSE_URL = "https://www.unicode.org/license.txt"; + +function arg(name, fallback) { + const index = process.argv.indexOf(`--${name}`); + return index === -1 ? fallback : process.argv[index + 1]; +} +const has = (name) => process.argv.includes(`--${name}`); + +const outPath = resolve(arg("out", DEFAULT_OUT)); +const limit = Number.parseInt(arg("limit", String(DEFAULT_LIMIT)), 10); +const checkOnly = has("check"); +const sourceArg = arg("source", null); + +/** Download Unihan.zip and extract Unihan_Readings.txt with the system unzip. */ +function fetchReadings() { + const dir = mkdtempSync(join(tmpdir(), "dst-unihan-")); + const zipPath = join(dir, "Unihan.zip"); + console.log(`downloading ${UNIHAN_URL}`); + const download = spawnSync("curl", ["-fsSL", "-o", zipPath, UNIHAN_URL], { encoding: "utf8" }); + if (download.status !== 0) { + throw new Error( + `cannot download Unihan.zip (${download.stderr?.trim() || "curl failed"}).\n` + + "Download it manually and pass --source .", + ); + } + const unzip = spawnSync("unzip", ["-o", "-q", zipPath, "Unihan_Readings.txt", "-d", dir], { + encoding: "utf8", + }); + if (unzip.status !== 0) { + throw new Error( + `cannot extract Unihan_Readings.txt (${unzip.stderr?.trim() || "unzip failed"}).\n` + + "Extract it manually and pass --source .", + ); + } + return { path: join(dir, "Unihan_Readings.txt"), cleanup: () => rmSync(dir, { recursive: true, force: true }) }; +} + +/** Parse Unihan_Readings.txt into per-character frequency and reading data. */ +export function parseUnihan(text) { + const sourceDate = /^#\s*Date:\s*(.+?)\s*$/m.exec(text)?.[1] ?? null; + const unicodeVersion = /^#\s*Unicode Version\s*(.+?)\s*$/m.exec(text)?.[1] ?? null; + + const frequencies = new Map(); + const readings = new Map(); + + for (const line of text.split("\n")) { + if (line.startsWith("#") || line.trim() === "") continue; + const [codeField, field, value] = line.split("\t"); + if (!codeField?.startsWith("U+") || !value) continue; + const character = String.fromCodePoint(Number.parseInt(codeField.slice(2), 16)); + + if (field === "kHanyuPinlu") { + let total = 0; + const entries = []; + for (const match of value.matchAll(/([A-Za-z\u00C0-\u024F\u0300-\u036F]+)\((\d+)\)/g)) { + const count = Number.parseInt(match[2], 10); + total += count; + entries.push({ reading: match[1], count }); + } + if (entries.length > 0) frequencies.set(character, { total, entries }); + } else if (field === "kMandarin") { + const primary = value.trim().split(/\s+/)[0]; + if (primary) readings.set(character, primary); + } + } + + return { sourceDate, unicodeVersion, frequencies, readings }; +} + +/** Build the asset object: top-N by frequency, code-point ordered, with provenance. */ +export function buildAsset({ sourceDate, unicodeVersion, frequencies, readings }, maxCharacters) { + const ranked = [...frequencies.entries()] + .map(([character, data]) => ({ character, total: data.total, entries: data.entries })) + .sort((left, right) => { + if (right.total !== left.total) return right.total - left.total; + return left.character.codePointAt(0) - right.character.codePointAt(0); + }); + + // Characters that carry a reading but no corpus frequency. They are real, common + // characters too — dropping them made the table cover 2404 of the 3500 it + // promised, and every one of the missing went to the explicit-slug fallback. + const unranked = [...readings.keys()] + .filter((character) => !frequencies.has(character)) + .sort((left, right) => left.codePointAt(0) - right.codePointAt(0)) + .map((character) => ({ character, total: 0, entries: [{ reading: readings.get(character), count: 0 }] })); + + const ordered = [...ranked, ...unranked]; + const selected = maxCharacters > 0 ? ordered.slice(0, maxCharacters) : ordered; + const characters = {}; + for (const entry of selected) { + const best = + readings.get(entry.character) ?? + [...entry.entries].sort((left, right) => right.count - left.count)[0].reading; + characters[entry.character] = best; + } + + return { + _comment: + "Derived pinyin table for slug generation. Generated file — run `node scripts/generate-pinyin.mjs` to refresh; `--check` fails on drift.", + source: { + database: "Unihan", + url: UNIHAN_URL, + file: "Unihan_Readings.txt", + fields: ["kMandarin", "kHanyuPinlu"], + unicode_version: unicodeVersion, + source_date: sourceDate, + }, + license: { + name: "Unicode License v3", + url: UNICODE_LICENSE_URL, + notice: + "Unihan data is Copyright © Unicode, Inc. and distributed under the Unicode License v3; see the URL above for the full text.", + }, + generated_by: "scripts/generate-pinyin.mjs", + selection: { + rule: "top N Han characters by summed kHanyuPinlu frequency; reading = first kMandarin, else the most frequent kHanyuPinlu reading", + limit: maxCharacters, + // Every character this table could have described, not just the frequency-ranked + // ones: "covered" is about the source, not about which slice we kept. + covered_characters: new Set([...frequencies.keys(), ...readings.keys()]).size, + }, + count: Object.keys(characters).length, + characters, + }; +} + +function main() { + let source = sourceArg; + let cleanup = () => {}; + if (!source) { + const fetched = fetchReadings(); + source = fetched.path; + cleanup = fetched.cleanup; + } else if (!existsSync(source)) { + throw new Error(`--source does not exist: ${source}`); + } + + const text = readFileSync(source, "utf8"); + const parsed = parseUnihan(text); + const asset = buildAsset(parsed, Number.isNaN(limit) ? DEFAULT_LIMIT : limit); + const serialized = `${JSON.stringify(asset, null, 2)}\n`; + cleanup(); + + console.log( + `unihan ${asset.source.unicode_version} (${asset.source.source_date}): ` + + `${asset.selection.covered_characters} covered characters, ` + + `${asset.count} written (limit ${asset.selection.limit})`, + ); + + if (checkOnly) { + if (!existsSync(outPath)) throw new Error(`asset missing: ${outPath}`); + const current = readFileSync(outPath, "utf8"); + if (current !== serialized) { + console.error(`DRIFT: ${outPath} does not match a fresh Unihan derivation.`); + console.error("Run `node scripts/generate-pinyin.mjs` and commit the result."); + process.exitCode = 1; + return; + } + console.log(`OK: ${outPath} matches the current Unihan derivation.`); + return; + } + + writeFileSync(outPath, serialized, "utf8"); + console.log(`wrote ${outPath} (${Buffer.byteLength(serialized)} bytes)`); +} + +if (process.argv[1] && resolve(process.argv[1]) === resolve(fileURLToPath(import.meta.url))) { + main(); +} diff --git a/tests/pinyin-slug.test.mjs b/tests/pinyin-slug.test.mjs new file mode 100644 index 00000000..3c081934 --- /dev/null +++ b/tests/pinyin-slug.test.mjs @@ -0,0 +1,131 @@ +/** + * The pinyin table and the slug discipline around it. + * + * `assets/pinyin.json` replaces the optional `pypinyin` dependency. The asset is + * generated by `scripts/generate-pinyin.mjs` from Unihan; here we check the + * shipped provenance, spot-check readings, and — most importantly — that a + * missing or incomplete table fails loudly instead of emitting a junk slug. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +import { + SlugResolutionError, + containsHan, + isHanCharacter, + loadPinyinTable, + pinyinSyllables, + readingToSyllable, + resetPinyinTable, + setSlugifyTable, +} from "../src/skill/slug.mjs"; +import { slugify } from "../src/skill/writer.mjs"; +import { buildAsset, parseUnihan } from "../scripts/generate-pinyin.mjs"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); +const assetPath = join(projectRoot, "assets", "pinyin.json"); + +test("the shipped asset carries its provenance and license", () => { + const asset = JSON.parse(readFileSync(assetPath, "utf8")); + assert.equal(asset.source.database, "Unihan"); + assert.equal(asset.source.file, "Unihan_Readings.txt"); + assert.match(asset.license.name, /Unicode License/); + assert.match(asset.license.url, /^https:\/\/www\.unicode\.org\//); + assert.equal(asset.selection.limit, 3500); + assert.equal(asset.count, Object.keys(asset.characters).length); + assert.ok(asset.count >= 3000, `expected ~3500 characters, got ${asset.count}`); + assert.ok(asset.selection.covered_characters >= asset.count); + assert.equal(asset.generated_by, "scripts/generate-pinyin.mjs"); +}); + +test("common characters map to their standard reading", () => { + const table = loadPinyinTable(); + assert.equal(table["周"], "zhōu"); + assert.equal(table["的"], "de"); + assert.equal(table["女"], "nǚ"); + assert.equal(table["张"], "zhāng"); +}); + +test("slugify matches pypinyin for Latin and covered Han names", () => { + assert.equal(slugify("Zadie Smith"), "zadie-smith"); + assert.equal(slugify("Élodie"), "elodie"); + assert.equal(slugify("A/B"), "a-b"); + assert.equal(slugify("周奇墨"), "zhou-qi-mo"); + assert.equal(slugify("徐志胜"), "xu-zhi-sheng"); + assert.equal(slugify("张伟"), "zhang-wei"); +}); + +test("ü becomes v, exactly like pypinyin", () => { + assert.equal(readingToSyllable("lǚ"), "lv"); + assert.equal(readingToSyllable("nǚ"), "nv"); + assert.equal(readingToSyllable("zhōu"), "zhou"); +}); + +test("an uncovered character asks for --slug instead of guessing", () => { + // 吕 is a common surname that the frequency-selected table does not cover. + assert.equal(loadPinyinTable()["吕"], undefined); + assert.throws( + () => slugify("吕丽"), + (error) => error instanceof SlugResolutionError && /--slug/.test(error.remedy), + ); +}); + +test("a missing table fails for Han names but still handles Latin ones", () => { + setSlugifyTable(null); + try { + assert.throws(() => slugify("周奇墨"), SlugResolutionError); + assert.equal(slugify("Zadie Smith"), "zadie-smith"); + assert.deepEqual(pinyinSyllables("A/B"), ["A/B"]); + } finally { + resetPinyinTable(); + } + assert.ok(loadPinyinTable()); +}); + +test("Han detection covers the CJK blocks the table is built from", () => { + assert.equal(isHanCharacter("周"), true); + assert.equal(isHanCharacter("A"), false); + assert.equal(containsHan("Zadie Smith"), false); + assert.equal(containsHan("周奇墨"), true); + assert.equal(containsHan("Élodie"), false); +}); + +const FIXTURE = [ + "# Unihan_Readings.txt", + "# Date: 2025-07-24 00:00:00 GMT [KL]", + "# Unicode Version 17.0.0", + "U+4E00\tkHanyuPinlu\tyī(32747)", + "U+4E00\tkMandarin\tyī", + "U+4E8C\tkHanyuPinlu\tèr(950)", + "U+4E8C\tkMandarin\tèr", + "U+4E09\tkHanyuPinlu\tsān(12) sàn(3)", + "U+4E09\tkMandarin\tsān sàn", + "U+56DB\tkHanyuPinlu\tsì(5)", + "", +].join("\n"); + +test("the generator is deterministic and prefers kMandarin", () => { + const parsed = parseUnihan(FIXTURE); + assert.equal(parsed.unicodeVersion, "17.0.0"); + assert.equal(parsed.frequencies.size, 4); + + const asset = buildAsset(parsed, 3); + assert.equal(asset.count, 3); + assert.equal(asset.selection.limit, 3); + // Code-point order, kMandarin (first reading) wins, ties break on frequency. + assert.deepEqual(Object.keys(asset.characters), ["一", "二", "三"]); + assert.equal(asset.characters["三"], "sān"); + // A character without kMandarin falls back to its most frequent reading. + const fallback = buildAsset({ ...parsed, readings: new Map() }, 4); + assert.equal(fallback.characters["三"], "sān"); + assert.equal(fallback.characters["四"], "sì"); + + // limit 0 keeps every covered character. + assert.equal(buildAsset(parsed, 0).count, 4); + // Same input, same bytes. + assert.equal(JSON.stringify(buildAsset(parsed, 3)), JSON.stringify(asset)); +}); From 4f8ca22dc44fe980afc1df77268c2a265563e410 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 10/15] =?UTF-8?q?chore(cli):=20=E9=87=8D=E5=BB=BA=202=20?= =?UTF-8?q?=E4=B8=AA=E6=96=87=E4=BB=B6=EF=BC=88=E5=8E=9F=E6=8F=90=E4=BA=A4?= =?UTF-8?q?=E4=BF=A1=E6=81=AF=E6=9C=AA=E8=AE=B0=E5=BD=95=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:(未记录) --- src/cli/paths.mjs | 36 ++++++++++++ tests/doctor-coverage.test.mjs | 100 +++++++++++++++++++++++++++++++++ 2 files changed, 136 insertions(+) create mode 100644 src/cli/paths.mjs create mode 100644 tests/doctor-coverage.test.mjs diff --git a/src/cli/paths.mjs b/src/cli/paths.mjs new file mode 100644 index 00000000..f93835ff --- /dev/null +++ b/src/cli/paths.mjs @@ -0,0 +1,36 @@ +/** + * Where a Skill family lives on disk. + * + * Two spellings of `--base-dir` grew apart: `harvest` / `parse-*` treat it as the + * directory that *contains* `skills/`, while `skill create` treats it as the + * storage root itself (`skills/colleague`). Both are reasonable, and neither is + * written down, so the same value silently means different places — a trap for + * users and for the coding agents that script this CLI. + * + * `resolveSkillsRoot` accepts both: the canonical `/skills/` wins + * whenever it exists (or `skills/` does), the storage-root reading is kept + * working for compatibility, and the caller gets a warning to surface when the + * legacy reading was used. + */ + +import { existsSync } from "node:fs"; +import { basename, join, resolve } from "node:path"; + +/** + * @param {{baseDir: string, family: string}} input + * @returns {{root: string, mode: "skills-root"|"storage-root", warning: string|null}} + */ +export function resolveSkillsRoot({ baseDir, family }) { + const base = resolve(baseDir); + const canonical = join(base, "skills", family); + if (existsSync(canonical)) return { root: canonical, mode: "skills-root", warning: null }; + if (basename(base) === family) return { root: base, mode: "storage-root", warning: null }; + if (existsSync(join(base, "skills"))) return { root: canonical, mode: "skills-root", warning: null }; + return { + root: base, + mode: "storage-root", + warning: + `--base-dir ${baseDir} has no skills/ directory, so it was read as the storage root itself (${base}). ` + + `The canonical spelling is --base-dir , i.e. ${join(base, "..", "..")} for this layout.`, + }; +} diff --git a/tests/doctor-coverage.test.mjs b/tests/doctor-coverage.test.mjs new file mode 100644 index 00000000..e5b46b5e --- /dev/null +++ b/tests/doctor-coverage.test.mjs @@ -0,0 +1,100 @@ +/** + * `doctor`'s evidence coverage and the two spellings of `--base-dir`. + * + * Doctor used to print a hardcoded `0 cited`, and to resolve `--base-dir` as + * `/` — one level short of the layout every other command writes, + * so it reported "no generated skills found" for any base directory. Both are + * pinned here: the count comes from `evidence/derived/*.json`, a dangling + * citation is named, and all three spellings of `--base-dir` land on the same + * skill without counting it three times. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +import { parseReceipt, runCli } from "./helpers/cli.mjs"; +import { resolveSkillsRoot } from "../src/cli/paths.mjs"; + +/** base/skills/colleague/ with a ledger and a derived claim. */ +function fixture({ dangling = false } = {}) { + const base = mkdtempSync(join(tmpdir(), "dst-doctor-")); + const skillDir = join(base, "skills", "colleague", "demo"); + mkdirSync(join(skillDir, "knowledge", "text"), { recursive: true }); + mkdirSync(join(skillDir, "evidence", "derived"), { recursive: true }); + writeFileSync( + join(skillDir, "meta.json"), + `${JSON.stringify({ slug: "demo", character: "colleague", display_name: "Demo", version: "v1" }, null, 2)}\n`, + "utf8", + ); + writeFileSync( + join(skillDir, "knowledge", "index.json"), + `${JSON.stringify([{ id: "k0001", kind: "subtitle", bytes: 10, sha256: "a".repeat(64), anchors: ["k0001", "k0002"] }], null, 2)}\n`, + "utf8", + ); + writeFileSync( + join(skillDir, "evidence", "derived", "voice.json"), + `${JSON.stringify({ kind: "voice", claims: [{ id: "voice.x", confidence: "high", evidence: dangling ? ["k0001", "k9999"] : ["k0001"], label: { zh: "x" }, value: 1 }] }, null, 2)}\n`, + "utf8", + ); + return { base, skillDir }; +} + +test("doctor counts the anchors the derivation cites, and names dangling ones", () => { + const clean = fixture(); + const dangling = fixture({ dangling: true }); + try { + const ok = parseReceipt(runCli(["doctor", "--base-dir", clean.base, "--json"]).stdout); + assert.equal(ok.skills, 1, "the skill must be found through --base-dir"); + assert.deepEqual(ok.anchors, { total: 2, cited: 1 }, "cited comes from evidence/derived, not a constant"); + assert.deepEqual(ok.warnings, []); + + const bad = parseReceipt(runCli(["doctor", "--base-dir", dangling.base, "--json"]).stdout); + assert.equal(bad.anchors.cited, 1); + assert.match(bad.warnings.join("\n"), /1 anchor\(s\) cited by evidence\/derived.*cannot be resolved/); + } finally { + rmSync(clean.base, { recursive: true, force: true }); + rmSync(dangling.base, { recursive: true, force: true }); + } +}); + +test("all three --base-dir spellings reach the same skill exactly once", () => { + const { base } = fixture(); + try { + const canonical = parseReceipt(runCli(["doctor", "--base-dir", base, "--json"]).stdout); + assert.equal(canonical.skills, 1); + + // The storage root itself (what `skill create --base-dir` writes). + const storage = parseReceipt(runCli(["doctor", "--base-dir", join(base, "skills", "colleague"), "--json"]).stdout); + assert.equal(storage.skills, 1, "a storage root must be inspected once, not once per family"); + assert.deepEqual(storage.anchors, canonical.anchors); + + // A bare directory: legacy reading, with a warning that says how to be explicit. + const bare = parseReceipt(runCli(["doctor", "--base-dir", join(base, "skills"), "--json"]).stdout); + assert.equal(bare.skills, 0); + assert.match(bare.warnings.join("\n"), /contains no skills\/ directory/); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); + +test("resolveSkillsRoot is deterministic about which spelling it took", () => { + const { base } = fixture(); + try { + const canonical = resolveSkillsRoot({ baseDir: base, family: "colleague" }); + assert.equal(canonical.mode, "skills-root"); + assert.equal(canonical.root, join(base, "skills", "colleague")); + + const storage = resolveSkillsRoot({ baseDir: join(base, "skills", "colleague"), family: "colleague" }); + assert.equal(storage.mode, "storage-root"); + assert.equal(storage.bare, false); + + const bare = resolveSkillsRoot({ baseDir: join(base, "skills"), family: "colleague" }); + assert.equal(bare.bare, true); + assert.match(bare.warning, /canonical spelling/); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); From a50c92f395fbc6be442df645c874f6d3e427c9aa Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 11/15] =?UTF-8?q?chore(evidence):=20=E9=87=8D=E5=BB=BA=202?= =?UTF-8?q?=20=E4=B8=AA=E6=96=87=E4=BB=B6=EF=BC=88=E5=8E=9F=E6=8F=90?= =?UTF-8?q?=E4=BA=A4=E4=BF=A1=E6=81=AF=E6=9C=AA=E8=AE=B0=E5=BD=95=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:(未记录) --- docs/evidence/pr-01-node-core.md | 53 +++++++++++++++++++++++ docs/evidence/pr-02-parse-zero-cred.md | 59 ++++++++++++++++++++++++++ 2 files changed, 112 insertions(+) create mode 100644 docs/evidence/pr-01-node-core.md create mode 100644 docs/evidence/pr-02-parse-zero-cred.md diff --git a/docs/evidence/pr-01-node-core.md b/docs/evidence/pr-01-node-core.md new file mode 100644 index 00000000..6797d69e --- /dev/null +++ b/docs/evidence/pr-01-node-core.md @@ -0,0 +1,53 @@ +# PR-01 · Node 单栈基座:入口 CLI + Skill 内核 + 安装器 + 拼音 + 测试移植 + +- 分支:`ds/01-node-core`(12 个提交,已本地合并进 `dot-skill-test`) +- 依赖:无。这一条是并行工作的基座:契约、命令注册表、验收脚本都由它落下来 +- 交付:38 个文件 / +11 050 行 + +## 1. 变更 + +| # | 提交 | 内容 | +| --- | --- | --- | +| 1 | `c383fca` | `src/skill/writer.mjs` + `src/skill/slug.mjs`:Python `skill_writer.py` 的 Node 移植,含拼音 slug | +| 2 | `3337c39` | `src/skill/versions.mjs`:版本归档(list / backup / rollback / cleanup),归档时间戳按 UTC 固定 | +| 3 | `579b159` | 归档列表确定性:同一天两次 `version list` 输出一致 | +| 4 | `f94675d` | `scripts/parity.mjs`:与迁移前 Python 的**逐字节 parity** 证据(按 rev 跑,不进日常门禁) | +| 5 | `441ceaf` | `src/commands/skill.mjs`:`skill create\|update\|list\|version` 接进注册表 | +| 6 | `9ebfa6e` | `src/install/hosts.mjs`:8 个宿主安装器合并成一个模块(路径矩阵单一出处) | +| 7 | `3436730` | 安装器测试移植到 `node --test`(claude / codex / openclaw / hermes) | +| 8 | `43c837e` | 已注册命令的 `--help` 打印双语两段 | +| 9 | `f6dcf87` | `listCommands()` 暴露命令名(供 doctor / prompt-lint / 审计共用) | +| 10 | `f97bc0c` | `install` / `uninstall` / `doctor` / `legacy` 适配器(旧 `python3 tools/*.py` 调用转发 + deprecation 警告) | +| 11 | `28b0c32` | `assets/pinyin.json`:从 Unihan 生成,去掉 `pypinyin` 运行时依赖 | +| 12 | `1da31aa` | 其余 Python 测试套件移植为 `node --test` | + +新增文件(节选):`src/commands/{index,skill,install,doctor,legacy}.mjs`、`src/cli/{args,receipt}.mjs`、 +`src/skill/{writer,presets,schema,slug,versions}.mjs`、`src/install/hosts.mjs`、`src/hosts/agents.mjs`、 +`assets/pinyin.json`、`scripts/{generate-pinyin,parity,acceptance}.mjs`、`docs/v2/{CONTRACT,ACCEPTANCE,STATUS}.md`、 +`tests/{dispatcher,commands,cli-lifecycle,skill-writer,pinyin-slug,install-*}.test.mjs`、 +公开语料夹具 `tests/fixtures/public-corpus/synthetic-interview/**`。 + +## 2. 验收(当前树,可复算) + +```bash +node --test tests/dispatcher.test.mjs tests/commands.test.mjs tests/cli-lifecycle.test.mjs \ + tests/skill-writer.test.mjs tests/pinyin-slug.test.mjs tests/install-*.test.mjs +# 33 个测试文件、330 个 test() 块:node --test tests/*.test.mjs → 340 pass / 0 fail +node bin/distilly.mjs --help # 22 个命令名,全部有中英两段 +node bin/distilly.mjs doctor # 宿主矩阵 8 个宿主,逐个报告是否已安装 +node scripts/parity.mjs # 历史 parity 证据(需要旧 rev) +``` + +要点:**零运行时依赖**(`package.json` 无 dependencies);入口唯一(`bin/distilly.mjs`); +`--json` 在任何命令上只输出一个对象(`tests/dispatcher.test.mjs` 断言); +命令注册表是唯一注册点,两段式命令名优先(`skill create` 赢过 `skill`)。 + +## 3. 回滚 + +- 逐提交可 revert;`src/commands/legacy.mjs` 单独 revert 会让旧 `tools/*.py` 调用直接报未知命令。 +- `assets/pinyin.json` 是生成物:`node scripts/generate-pinyin.mjs` 可重现(`--check` 防漂移)。 + +## 4. 已知缺口 + +- `scripts/parity.mjs` 需要一份迁移前的 rev 才能跑:parity 是历史证据,不是日常门禁。 +- 这一条只交付 skill/install/doctor 内核;`harvest` / `parse-*` / `view` / `collect` 由 #02/#03/#07 交付。 diff --git a/docs/evidence/pr-02-parse-zero-cred.md b/docs/evidence/pr-02-parse-zero-cred.md new file mode 100644 index 00000000..b962cd52 --- /dev/null +++ b/docs/evidence/pr-02-parse-zero-cred.md @@ -0,0 +1,59 @@ +# PR-02 · 零凭据解析:`knowledge/` 账本 + 锚点 + chat/subtitle/archive 解析 + +- 分支:`ds/02-parse-zero-cred`(6 个提交,已本地合并进 `dot-skill-test`) +- 依赖:契约(`docs/v2/CONTRACT.md`,由 #01 冻结) +- 交付:31 个文件 / +5 905 行 + +## 1. 变更 + +| # | 提交 | 内容 | +| --- | --- | --- | +| 1 | `05ff594` | 冻结 v2 契约:磁盘布局、回执形状、密钥纪律、computer-use 同意 | +| 2 | `6f95373` | `src/knowledge/store.mjs`:`knowledge/raw/` 字节保险库(逐字落盘 + 读回校验 + 路径逃逸拒绝) | +| 3 | `66bc96f` | `src/knowledge/{anchors,ledger}.mjs`:段落锚点分配、只增账本、`units`/`anchor_detail`、去重(sha256 + origin) | +| 4 | `34e20a3` | `src/parse/subtitle.mjs`:`.srt`/`.vtt` → 一条 cue 一个锚点单元(含说话人、时间码、字节区间) | +| 5 | `1fd53af` | `src/parse/common.mjs`:零依赖读共享 zip 容器(中央目录、CRC 校验、成员流式解压) | +| 6 | `6c42d88` | `src/parse/chat.mjs`:ChatGPT / Claude / Slack / Telegram / Discord / Instagram 导出,其余按名拒绝 | + +新增文件:`src/knowledge/{store,anchors,ledger}.mjs`、`src/parse/{common,chat,subtitle}.mjs`、 +`tests/{knowledge-store,knowledge-anchors,knowledge-ledger,parse-chat,parse-subtitle}.test.mjs`、 +`tests/fixtures/parse/{chat,subtitle}/**`、`.gitattributes`(夹具按字节保真,禁换行转换)。 + +## 2. 磁盘契约(这一条定下来,后面所有分支都按它写) + +``` +skills///knowledge/ + raw//… 原样字节,只增不改 + text/.md 归一化正文,段落锚点 [k0012] / [k0012:t3] + index.json 账本:{id,kind,origin,fetched_at,bytes,sha256,credentialed,method,warnings[]} +``` + +- **锚点必须能回指**:`resolveLedgerAnchor(ledger, anchor)` 返回文本与字节区间; + 容器里读出来的文本(OOXML 成员、去标签的 HTML)标 `synthetic` 且**报 null 字节区间**,不伪造偏移。 +- **幂等**:同一份字节(sha256 相同且 origin 相同)重复导入不新增条目。 +- **响亮拒绝**:不认识的格式按文件名进 warnings,不静默跳过。 + +## 3. 验收(当前树,可复算) + +```bash +node --test tests/knowledge-store.test.mjs tests/knowledge-anchors.test.mjs \ + tests/knowledge-ledger.test.mjs tests/parse-chat.test.mjs tests/parse-subtitle.test.mjs +node scripts/acceptance.mjs # harvest / 账本 / 幂等 / 锚点回指四个阶段 +``` + +`scripts/acceptance.mjs` 里与本条直接相关的断言:回执形状(sha256 + 字节数)、 +重复 harvest 幂等、账本里有锚点、派生结论锚点全部可回指。 + +## 4. 后续修复(合并后由集成轮补上,见后续 PR 文档) + +- `assignAnchorsToText` 曾把段落渲染成 `k0012 text`(无方括号),而本条的读者要求 + `[k0012] text`:派生层因此恒为空。修法与前后对比见 + `docs/evidence/pr-09-evidence-spine-blind-test.md`。 +- 归一化正文曾丢掉说话人与时间(本条的 parser 有元数据,但没进正文):修法与前后对比见 + `docs/evidence/pr-11-attribution.md`。 + +## 5. 回滚 + +- `src/knowledge/**` 与 `src/parse/**` 是后面所有解析命令的地基:回滚本条会让 + `harvest` / `parse-*` / `retrospect` 全部失效,需回滚到 #01 的 CLI 骨架状态。 +- `.gitattributes` 的字节保真规则不要单独 revert:夹具在 CRLF 平台上会被改写,测试随之漂移。 From 164c791e318debdcbd27a5242c748d1d84e9f653 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 12/15] =?UTF-8?q?chore(workflows):=20=E9=87=8D=E5=BB=BA=20?= =?UTF-8?q?.github/workflows/ci.yml=EF=BC=88=E5=8E=9F=E6=8F=90=E4=BA=A4?= =?UTF-8?q?=E4=BF=A1=E6=81=AF=E6=9C=AA=E8=AE=B0=E5=BD=95=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:(未记录) --- .github/workflows/ci.yml | 88 +++++++++++++++++++++++++++------------- 1 file changed, 59 insertions(+), 29 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index a7d5565f..da8d4dbb 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -2,53 +2,83 @@ name: CI on: push: - branches: [dot-skill, main] + branches: [dot-skill-test, dot-skill, main] pull_request: - branches: [dot-skill, main] + branches: [dot-skill-test, dot-skill, main] jobs: test: - name: Python ${{ matrix.python-version }} + name: Node ${{ matrix.node-version }} runs-on: ubuntu-latest strategy: fail-fast: false matrix: - python-version: ["3.9", "3.11"] + node-version: ["20", "22"] steps: - uses: actions/checkout@v4 - - name: Set up Python ${{ matrix.python-version }} - uses: actions/setup-python@v5 + - uses: actions/setup-node@v4 with: - python-version: ${{ matrix.python-version }} - cache: pip + node-version: ${{ matrix.node-version }} - - name: Install dependencies + - name: Syntax-check every module run: | - python -m pip install --upgrade pip - if [ -f requirements.txt ]; then pip install -r requirements.txt; fi + find bin src scripts tests -name '*.mjs' -print0 | xargs -0 -n1 node --check - - name: Compile all Python sources - run: python -m compileall -q tools + - name: Unit tests + run: node --test - - name: Run unit tests - run: | - if [ -d tests ]; then - python -m unittest discover -s tests -p 'test_*.py' -v - else - echo "No tests/ directory yet — skipping unittest discover." - fi - - lint: - name: Ruff + - name: Prompt contract lint + run: node scripts/prompt-lint.mjs + + - name: Skill template freshness + run: node scripts/generate-template.mjs --check + + # Every demand of the v2 objective, mapped to the artefact that proves it. + # Acceptance itself runs in the next job, so this is scope-only. + - name: Objective audit + run: node scripts/audit-objective.mjs --skip-acceptance + + # Release hygiene: versions, schema marker, carried directories, gates. + - name: Release check + run: node scripts/check_release.mjs + + acceptance: + name: Acceptance (public corpus) runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - - uses: actions/setup-python@v5 + + - uses: actions/setup-node@v4 + with: + node-version: "22" + + # The visual-check phase drives a real browser; playwright is a dev + # dependency and never ships with the package. + - name: Install Playwright + run: | + npm install --no-save playwright@1.62.1 + npx playwright install --with-deps chromium + + - name: End-to-end acceptance on the bundled corpus + run: node scripts/acceptance.mjs --evidence "$RUNNER_TEMP/evidence" + + # The second public corpus is multi-source (two chat exports + a document) + # and dated, so the same phases run on a corpus the derivation was designed for. + - name: End-to-end acceptance on the multi-source corpus + run: node scripts/acceptance.mjs --corpus tests/fixtures/public-corpus/synthetic-multisource --evidence "$RUNNER_TEMP/evidence-multisource" + + # Scope audit *with* the acceptance run it normally performs. + - name: Objective audit (includes acceptance) + run: DISTILLY_PLAYWRIGHT_ROOT="$PWD" node scripts/audit-objective.mjs + + - uses: actions/upload-artifact@v4 + if: always() with: - python-version: "3.11" - - name: Install ruff - run: pip install ruff - - name: Run ruff (non-blocking for now) - run: ruff check tools/ || true + name: acceptance-evidence + path: | + ${{ runner.temp }}/evidence + ${{ runner.temp }}/evidence-multisource + if-no-files-found: ignore + From 4932aa6d784179e3231fc5f05b0ccabfa47fca94 Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:58 +0800 Subject: [PATCH 13/15] =?UTF-8?q?chore(INSTALL.md):=20=E9=87=8D=E5=BB=BA?= =?UTF-8?q?=2023=20=E4=B8=AA=E6=96=87=E4=BB=B6=EF=BC=88=E5=8E=9F=E6=8F=90?= =?UTF-8?q?=E4=BA=A4=E4=BF=A1=E6=81=AF=E6=9C=AA=E8=AE=B0=E5=BD=95=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:本提交由会话转录重建,提交信息取自原分支(ds/01-node-core)。 文件内容为集成分支上的最终态,不是当时那一刻的中间态——原分支的 per-commit 文件树随 /tmp 清空丢失,转录只保留了提交信息与 git add 的路径清单。 原提交信息:(未记录) --- INSTALL.md | 37 ++ INSTALL_EN.md | 46 ++ package.json | 9 +- scripts/check_release.mjs | 124 ++++ src/cli/args.mjs | 103 ++++ src/cli/entry.mjs | 33 + src/cli/receipt.mjs | 118 ++++ src/commands/migrate.mjs | 110 ++++ src/commands/skill.mjs | 480 +++++++++++++++ src/skill/migrate.mjs | 118 ++++ tests/cli-lifecycle.test.mjs | 464 ++++++++++++++ tests/commands.test.mjs | 235 +++++++ tests/entry-point.test.mjs | 87 +++ tests/entrypoint-docs.test.mjs | 88 +++ tests/helpers/cli.mjs | 27 + tests/install-claude-generated-skill.test.mjs | 84 +++ tests/install-generated-skill.test.mjs | 135 ++++ tests/install-hermes-skill.test.mjs | 116 ++++ tests/install-openclaw-and-codex.test.mjs | 157 +++++ tests/package-payload.test.mjs | 147 +++++ tests/release-check.test.mjs | 37 ++ tests/schema-migration.test.mjs | 165 +++++ tests/skill-writer.test.mjs | 574 ++++++++++++++++++ 23 files changed, 3491 insertions(+), 3 deletions(-) create mode 100644 scripts/check_release.mjs create mode 100644 src/cli/args.mjs create mode 100644 src/cli/entry.mjs create mode 100644 src/cli/receipt.mjs create mode 100644 src/commands/migrate.mjs create mode 100644 src/commands/skill.mjs create mode 100644 src/skill/migrate.mjs create mode 100644 tests/cli-lifecycle.test.mjs create mode 100644 tests/commands.test.mjs create mode 100644 tests/entry-point.test.mjs create mode 100644 tests/entrypoint-docs.test.mjs create mode 100644 tests/helpers/cli.mjs create mode 100644 tests/install-claude-generated-skill.test.mjs create mode 100644 tests/install-generated-skill.test.mjs create mode 100644 tests/install-hermes-skill.test.mjs create mode 100644 tests/install-openclaw-and-codex.test.mjs create mode 100644 tests/package-payload.test.mjs create mode 100644 tests/release-check.test.mjs create mode 100644 tests/schema-migration.test.mjs create mode 100644 tests/skill-writer.test.mjs diff --git a/INSTALL.md b/INSTALL.md index f051a40c..09e3b034 100644 --- a/INSTALL.md +++ b/INSTALL.md @@ -25,8 +25,44 @@ --- + + +## v2:命令入口与宿主适配 + +v2 只有一个命令入口:**`bin/distilly.mjs`**(Node ≥ 20,零依赖,见 `docs/v2/CONTRACT.md`)。 + +```bash +node bin/distilly.mjs install # 装到该宿主的全局 Skill 目录 +node bin/distilly.mjs install --force # 已有安装先备份成 *.backup-<时间戳> 再替换 +node bin/distilly.mjs install --path

# 装到自定义路径(末段目录必须叫 distilly) +node bin/distilly.mjs --help +``` + +- **宿主 id、全局/项目级目录、确切安装命令、双语注意事项、装完怎么验证**:见 + **[docs/v2/HOSTS.md](docs/v2/HOSTS.md)**。该表由 `src/hosts/agents.mjs` 生成,`tests/agents.test.mjs` + 强制它与 `bin/distilly.mjs` 的落盘目录一致。 +- 当前支持 8 个宿主:`claude-code` · `codex` · `opencode` · `openclaw` · `hermes` · `deepseek-harness` · + `grok-build` · `pi`。别名:`claude`、`deepseek`、`grok`。 +- 两条路线等价:`npx -y skills add titanwings/distilly --skill distilly …`(AgentSkills CLI)或直接 + `git clone https://github.com/titanwings/distilly <目标目录>`;逐字命令同样在 `docs/v2/HOSTS.md`。 + + + +### ⚠️ Deprecated:`python3 tools/*.py` 安装器 + +下面「选择你的平台」各节里的 `python3 tools/install_*_skill.py` 与手工 `git clone` 是**迁移期兼容路径,已废弃**: +v2 不再要求用户手动跑 Python。契约(`docs/v2/CONTRACT.md` §1)约定旧的 `python3 tools/xxx.py` 调用由 +`bin/distilly.mjs` 转发并打印 deprecation 警告,转发层在 PR③ 删除;在当前集成分支上这些命令仍然等价于 +直接执行对应的 Python 脚本。旧内容只为排查老安装而保留,**新安装请走 `bin/distilly.mjs` 或 +`docs/v2/HOSTS.md` 里的一行命令**。 + +--- + ## 选择你的平台 +> ⚠️ **Deprecated(旧安装路径)**:本节保留旧版按平台展开的说明。宿主目录与确切命令的最新版本在 +> **[docs/v2/HOSTS.md](docs/v2/HOSTS.md)**;下面的 `python3 tools/*.py` 调用见上一节的废弃说明。 + ### A. Claude Code(推荐) 本项目遵循官方 [AgentSkills](https://agentskills.io) 标准,整个 repo 就是 skill 目录。克隆到 Claude skills 目录即可: @@ -533,3 +569,4 @@ distilly/ ← clone 到宿主的 skills/distilly/(例如 .claude ├── versions/ # 历史版本 └── knowledge/ # 原始材料归档 ``` + diff --git a/INSTALL_EN.md b/INSTALL_EN.md index b916f542..fb91f756 100644 --- a/INSTALL_EN.md +++ b/INSTALL_EN.md @@ -3,8 +3,53 @@ > Distilly was formerly known as **Colleague Skill / colleague-skill**. The creator > name and canonical install directory are now `distilly`. + + +## v2: command entrypoint and host adaptation + +v2 has exactly one command entrypoint: **`bin/distilly.mjs`** (Node >= 20, zero +dependencies, see `docs/v2/CONTRACT.md`). + +```bash +node bin/distilly.mjs install # install into that host's global Skill directory +node bin/distilly.mjs install --force # back up an existing install as *.backup-, then replace +node bin/distilly.mjs install --path

# install into a custom path (final directory must be `distilly`) +node bin/distilly.mjs --help +``` + +- **Host ids, global/project directories, the exact install commands, bilingual + notes and how to verify an install** live in + **[docs/v2/HOSTS.md](docs/v2/HOSTS.md)**. That table is generated from + `src/hosts/agents.mjs`, and `tests/agents.test.mjs` forces it to agree with the + destinations in `bin/distilly.mjs`. +- Eight hosts are supported today: `claude-code`, `codex`, `opencode`, + `openclaw`, `hermes`, `deepseek-harness`, `grok-build`, `pi`. Aliases: + `claude`, `deepseek`, `grok`. +- The two routes are equivalent — `npx -y skills add titanwings/distilly + --skill distilly …` (AgentSkills CLI) or a plain + `git clone https://github.com/titanwings/distilly `; both are quoted + verbatim in `docs/v2/HOSTS.md`. + + + +### ⚠️ Deprecated: the `python3 tools/*.py` installers + +The `python3 tools/install_*_skill.py` calls and manual clones below are +**migration-era compatibility paths and are deprecated**: v2 no longer asks users +to run Python by hand. The contract (`docs/v2/CONTRACT.md` §1) says the entrypoint +forwards the old `python3 tools/xxx.py` calls with a deprecation warning and that +the forwarding layer is removed in PR③; on the current integration branch those +commands are still equivalent to running the Python script directly. The old +sections are kept for troubleshooting legacy installs only — **use +`bin/distilly.mjs` or the one-liners in `docs/v2/HOSTS.md` for new installs**. + ## Install Distilly +> ⚠️ **Deprecated (legacy install path)**: this section keeps the old per-host +> walkthrough. The current host directories and exact commands are in +> **[docs/v2/HOSTS.md](docs/v2/HOSTS.md)**; the `python3 tools/*.py` calls are +> explained in the deprecation note above. + Clone the repository into a Skills directory discovered by your host, keeping the destination directory name `distilly`: @@ -157,3 +202,4 @@ keep only copyright-safe paraphrases with source URLs in research notes, and delete the temporary file after review. Xquik is independent of X Corp. “Twitter” and “X” are trademarks of X Corp. + diff --git a/package.json b/package.json index 453f911c..81fefabe 100644 --- a/package.json +++ b/package.json @@ -8,17 +8,19 @@ }, "files": [ "bin/", + "src/", + "assets/", + "scripts/", "SKILL.md", "prompts/", "references/", - "tools/", - "requirements.txt", "INSTALL.md", "INSTALL_EN.md", "LICENSE", "CITATION.cff" ], "scripts": { + "test": "node --test \"tests/*.test.mjs\"", "prepack": "node bin/distilly.mjs --check-package" }, "keywords": [ @@ -41,9 +43,10 @@ "url": "https://github.com/titanwings/distilly/issues" }, "engines": { - "node": ">=18" + "node": ">=20" }, "publishConfig": { "registry": "https://npm.pkg.github.com" } } + diff --git a/scripts/check_release.mjs b/scripts/check_release.mjs new file mode 100644 index 00000000..7ac374c4 --- /dev/null +++ b/scripts/check_release.mjs @@ -0,0 +1,124 @@ +#!/usr/bin/env node +/** + * Release readiness for this package: the things a version bump must not forget. + * + * `acceptance.mjs` proves the pipeline works and `audit-objective.mjs` proves the + * scope is closed; neither notices that the version printed by `--version` drifted + * from `package.json`, that a migrated ledger lost its schema marker, or that a + * `.py` file came back. Those are release-time questions, so they live here. + * + * node scripts/check_release.mjs [--json] [--tag vX.Y.Z] + * + * Exit code 1 when any check fails. `--tag` additionally asserts that the release + * tag names the same version the package declares. + */ + +import { execFileSync } from "node:child_process"; +import { existsSync, readFileSync } from "node:fs"; +import { join, resolve } from "node:path"; + +import { SCHEMA_VERSION } from "../src/skill/schema.mjs"; +import { LEDGER_SCHEMA_VERSION } from "../src/knowledge/ledger.mjs"; + +const root = resolve(import.meta.dirname, ".."); +const json = process.argv.includes("--json"); +const tagIndex = process.argv.indexOf("--tag"); +const tag = tagIndex === -1 ? null : process.argv[tagIndex + 1]; + +const rows = []; +const record = (name, ok, evidence) => rows.push({ name, ok: Boolean(ok), evidence }); + +const read = (relative) => readFileSync(join(root, relative), "utf8"); +const pkg = JSON.parse(read("package.json")); + +/* 1 — one version, everywhere it is printed ---------------------------------- */ +{ + const binVersion = execFileSync(process.execPath, [join(root, "bin", "distilly.mjs"), "--version"], { encoding: "utf8" }).trim(); + const skill = existsSync(join(root, "SKILL.md")) ? read("SKILL.md") : ""; + const skillVersion = /^version:\s*"?([^"\n]+)"?/m.exec(skill)?.[1]?.trim() ?? null; + const consistent = binVersion === pkg.version && (skillVersion === null || skillVersion === pkg.version); + record( + "版本一致(package.json / --version / SKILL.md)", + consistent, + `package.json ${pkg.version}, --version ${binVersion}, SKILL.md ${skillVersion ?? "(未声明)"}`, + ); + if (tag !== null) { + record("发布 tag 指向同一版本", tag === `v${pkg.version}` || tag === pkg.version, `--tag ${tag} vs ${pkg.version}`); + } +} + +/* 2 — schemas are the frozen ones, and a migration exists -------------------- */ +{ + const migrate = existsSync(join(root, "src", "skill", "migrate.mjs")); + const migrationTest = read(join("tests", "schema-migration.test.mjs")); + record( + `schema v${SCHEMA_VERSION} 与账本 v${LEDGER_SCHEMA_VERSION},迁移脚本在`, + SCHEMA_VERSION === "4" && migrate && /idempot/i.test(migrationTest), + `SCHEMA_VERSION=${SCHEMA_VERSION}, LEDGER_SCHEMA_VERSION=${LEDGER_SCHEMA_VERSION}, src/skill/migrate.mjs=${migrate}, 幂等断言=${/idempot/i.test(migrationTest)}`, + ); +} + +/* 3 — the installers still carry the evidence spine -------------------------- */ +{ + const hosts = read(join("src", "install", "hosts.mjs")); + const carried = ["knowledge/raw", "knowledge/text", "evidence", "views"].every((name) => hosts.includes(name)); + record( + "安装器携带 evidence spine(knowledge/raw、knowledge/text、evidence、views)", + carried, + carried ? "CARRIED_DIRECTORIES 覆盖四项" : "CARRIED_DIRECTORIES 缺项", + ); +} + +/* 4 — the package is zero-dependency and Python-free ------------------------- */ +{ + const dependencies = Object.keys(pkg.dependencies ?? {}); + const tracked = execFileSync("git", ["ls-files"], { cwd: root, encoding: "utf8" }).split("\n"); + const python = tracked.filter((path) => /\.py$/.test(path) || /(^|\/)requirements\.txt$/.test(path)); + record( + "零运行时依赖且没有 Python 残留", + dependencies.length === 0 && python.length === 0, + `dependencies: ${dependencies.length === 0 ? "none" : dependencies.join(", ")}; tracked .py / requirements.txt: ${python.length}`, + ); +} + +/* 5 — generated artefacts are in sync with their sources -------------------- */ +{ + const template = execFileSync(process.execPath, [join(root, "scripts", "generate-template.mjs"), "--check"], { encoding: "utf8" }).trim(); + const pinyin = existsSync(join(root, "assets", "pinyin.json")); + record( + "生成物与源同步(模板 / 拼音表)", + /up to date/.test(template) && pinyin, + `${template.split("\n").pop()}; assets/pinyin.json: ${pinyin}`, + ); +} + +/* 6 — the gates a release claims are runnable -------------------------------- */ +{ + const gates = ["scripts/acceptance.mjs", "scripts/audit-objective.mjs", "scripts/prompt-lint.mjs", "scripts/visual-check.mjs", "scripts/split-corpus.mjs", "scripts/blind-test.mjs"]; + const missing = gates.filter((path) => !existsSync(join(root, path))); + const ci = read(join(".github", "workflows", "ci.yml")); + const wired = ["node --test", "acceptance.mjs", "prompt-lint.mjs", "audit-objective.mjs"].every((needle) => ci.includes(needle)); + record( + "发布所依赖的门禁都在,且 CI 会跑", + missing.length === 0 && wired, + `missing: ${missing.length === 0 ? "none" : missing.join(", ")}; CI 覆盖 node --test / acceptance / prompt-lint / audit: ${wired}`, + ); +} + +/* 7 — documentation a release points at ------------------------------------- */ +{ + const docs = ["docs/v2/CONTRACT.md", "docs/v2/ACCEPTANCE.md", "docs/v2/STATUS.md", "docs/v2/MIGRATION.md", "docs/v2/IDENTITY.md", "README.md"]; + const missing = docs.filter((path) => !existsSync(join(root, path))); + record("发布指向的文档都在", missing.length === 0, missing.length === 0 ? `${docs.length} 份文档就位` : `缺: ${missing.join(", ")}`); +} + +const failed = rows.filter((row) => !row.ok); +if (json) { + console.log(JSON.stringify({ ok: failed.length === 0, version: pkg.version, schema: SCHEMA_VERSION, rows }, null, 2)); +} else { + console.log("发布检查 / release check\n"); + for (const row of rows) console.log(`${row.ok ? "✅" : "❌"} ${row.name}\n ${row.evidence}`); + console.log(`\n${rows.length - failed.length}/${rows.length} 项通过${failed.length === 0 ? "" : `;未通过:${failed.map((row) => row.name).join(";")}`}`); +} +process.exit(failed.length === 0 ? 0 : 1); + diff --git a/src/cli/args.mjs b/src/cli/args.mjs new file mode 100644 index 00000000..30a4da19 --- /dev/null +++ b/src/cli/args.mjs @@ -0,0 +1,103 @@ +/** + * Minimal argument parser for the Distilly CLI (zero dependencies). + * + * Mirrors the subset of Python's argparse behaviour the ported tools relied on: + * `--flag value`, `--flag=value`, boolean switches, repeated options and + * positionals. Unknown options are a hard error — the CLI never guesses. + */ + +/** Raised for user-facing argument problems (exit code 1). */ +export class ArgError extends Error { + constructor(message) { + super(message); + this.name = "ArgError"; + } +} + +/** + * @typedef {object} OptionSpec + * @property {'boolean'|'string'} type + * @property {string} [alias] short flag without dashes, e.g. `o` + * @property {string} [help] + * @property {boolean} [multiple] collect repeats into an array + * @property {string} [value] metavar shown in help + */ + +function optionNames(longName, spec) { + const names = [`--${longName}`]; + if (spec.alias) names.push(`-${spec.alias}`); + return names; +} + +/** + * @param {string[]} argv + * @param {Record} spec + * @returns {{flags: Record, positionals: string[]}} + */ +export function parseArgs(argv, spec = {}) { + const byName = new Map(); + for (const [longName, option] of Object.entries(spec)) { + for (const name of optionNames(longName, option)) byName.set(name, longName); + } + + const flags = {}; + for (const [longName, option] of Object.entries(spec)) { + if (option.multiple) flags[longName] = []; + else if (option.type === "boolean") flags[longName] = false; + else flags[longName] = undefined; + } + + const positionals = []; + let index = 0; + while (index < argv.length) { + const arg = argv[index]; + if (arg === "--") { + positionals.push(...argv.slice(index + 1)); + break; + } + if (arg.startsWith("-") && arg !== "-") { + const equals = arg.indexOf("="); + const name = equals === -1 ? arg : arg.slice(0, equals); + const inlineValue = equals === -1 ? undefined : arg.slice(equals + 1); + const longName = byName.get(name); + if (!longName) throw new ArgError(`unrecognized argument: ${name}`); + const option = spec[longName]; + if (option.type === "boolean") { + if (inlineValue !== undefined) { + throw new ArgError(`${name} does not take a value`); + } + flags[longName] = true; + } else { + const value = inlineValue !== undefined ? inlineValue : argv[index + 1]; + if (value === undefined || (inlineValue === undefined && value.startsWith("-") && value !== "-")) { + throw new ArgError(`${name} requires a value`); + } + if (inlineValue === undefined) index += 1; + if (option.multiple) flags[longName].push(value); + else flags[longName] = value; + } + index += 1; + continue; + } + positionals.push(arg); + index += 1; + } + + return { flags, positionals }; +} + +/** True when the argument list asks for help. */ +export function wantsHelp(argv) { + return argv.includes("--help") || argv.includes("-h"); +} + +/** Render ` ` fragments for a usage line. */ +export function usageFragment(spec = {}) { + const parts = []; + for (const [longName, option] of Object.entries(spec)) { + if (option.hidden) continue; + const bare = option.alias ? `-${option.alias}, --${longName}` : `--${longName}`; + parts.push(option.type === "boolean" ? `[${bare}]` : `[${bare} <${option.value ?? "value"}>]`); + } + return parts.join(" "); +} diff --git a/src/cli/entry.mjs b/src/cli/entry.mjs new file mode 100644 index 00000000..24dff7bd --- /dev/null +++ b/src/cli/entry.mjs @@ -0,0 +1,33 @@ +/** + * "Am I the process entry point?" — the guard every executable here needs. + * + * `bin/distilly.mjs` and the two runnable scripts under `scripts/` must dispatch + * only when they were invoked directly, because tests import their internals. + * The obvious spelling is wrong in a way that fails *silently*: + * + * resolve(process.argv[1]) === fileURLToPath(import.meta.url) + * + * `import.meta.url` is always the realpath, while `argv[1]` is whatever the + * caller wrote. Under a symlink the two differ, the guard concludes "I was + * imported", and the program exits 0 having done nothing. That is not an edge + * case: `/tmp` is a symlink to `/private/tmp` on macOS, and — more importantly — + * an npm `bin` shim in `node_modules/.bin/` is a symlink, so a published + * `distilly` would have silently ignored every command. + * + * Resolving both sides is the fix. A path that cannot be resolved (an eval, a + * repl) is not an entry point, which is also the honest answer. + */ + +import { realpathSync } from "node:fs"; +import { fileURLToPath } from "node:url"; + +/** True when `moduleUrl` names the file the process was started with. */ +export function isEntryPoint(moduleUrl) { + const invoked = process.argv[1]; + if (invoked === undefined || invoked === "") return false; + try { + return realpathSync(invoked) === realpathSync(fileURLToPath(moduleUrl)); + } catch { + return false; + } +} diff --git a/src/cli/receipt.mjs b/src/cli/receipt.mjs new file mode 100644 index 00000000..aea2d7a6 --- /dev/null +++ b/src/cli/receipt.mjs @@ -0,0 +1,118 @@ +/** + * CLI receipts, output routing and file fingerprints. + * + * The receipt shape is frozen by `docs/v2/CONTRACT.md` §3: + * `{command, person, ok, inputs, outputs, anchors, warnings, unavailable}`. + * Every entry in `inputs`/`outputs` carries `{path, sha256, bytes}`. + * + * `--json` writes the receipt to stdout as the only stdout content (human text + * moves to stderr) so `JSON.parse(stdout)` always succeeds. + */ + +import { createHash } from "node:crypto"; +import { readFileSync, statSync } from "node:fs"; +import { relative, resolve, sep } from "node:path"; + +/** User-facing failure with an explicit remedy. */ +export class CliError extends Error { + constructor(message, { code = "error", remedy = "", exitCode = 1 } = {}) { + super(message); + this.name = "CliError"; + this.code = code; + this.remedy = remedy; + this.exitCode = exitCode; + } +} + +export function sha256Buffer(buffer) { + return createHash("sha256").update(buffer).digest("hex"); +} + +export function sha256Text(text) { + return sha256Buffer(Buffer.from(text, "utf8")); +} + +/** `{path, sha256, bytes}` for a file, or null when it is missing. */ +export function describeFile(filePath, { cwd = process.cwd() } = {}) { + let buffer; + try { + buffer = readFileSync(filePath); + } catch { + return null; + } + return { + path: displayPath(filePath, cwd), + sha256: sha256Buffer(buffer), + bytes: buffer.length, + }; +} + +/** `{path, bytes}` without hashing (used for directories and removals). */ +export function describePath(filePath, { cwd = process.cwd() } = {}) { + return { path: displayPath(filePath, cwd), bytes: directoryBytes(filePath) }; +} + +export function directoryBytes(dirPath) { + try { + return statSync(dirPath).size; + } catch { + return 0; + } +} + +/** Paths inside the working directory are reported relatively, like the old CLI. */ +export function displayPath(filePath, cwd = process.cwd()) { + const absolute = resolve(filePath); + const rel = relative(resolve(cwd), absolute); + if (rel === "") return "."; + if (!rel.startsWith("..") && !rel.startsWith(`${sep}..`)) return rel; + return absolute; +} + +/** Build the contract receipt object (field order is contractual). */ +export function createReceipt(command, options = {}) { + const receipt = { + command, + person: options.person ?? null, + ok: options.ok ?? true, + inputs: options.inputs ?? [], + outputs: options.outputs ?? [], + anchors: options.anchors ?? { total: 0, cited: 0 }, + warnings: options.warnings ?? [], + unavailable: options.unavailable ?? [], + }; + if (options.error) receipt.error = options.error; + return receipt; +} + +/** + * Route human text and machine receipts. + * In `--json` mode stdout carries exactly one JSON object; prose goes to stderr. + */ +export function createReporter(json, { stdout = process.stdout, stderr = process.stderr } = {}) { + const lines = []; + return { + json, + line(text) { + lines.push(text); + if (json) stderr.write(`${text}\n`); + else stdout.write(`${text}\n`); + }, + warn(text) { + stderr.write(`${text}\n`); + }, + /** + * A diagnostic that always goes to stderr, in both modes. Command output is + * `line()`; this is for "why the command failed", which must never pollute + * the machine channel — and must still be visible when `--json` is set. + */ + error(text) { + stderr.write(`${text}\n`); + }, + /** Write the receipt last so `JSON.parse(stdout)` sees a single object. */ + finish(receipt) { + if (!json) return; + stdout.write(`${JSON.stringify(receipt, null, 2)}\n`); + }, + }; +} diff --git a/src/commands/migrate.mjs b/src/commands/migrate.mjs new file mode 100644 index 00000000..88ece416 --- /dev/null +++ b/src/commands/migrate.mjs @@ -0,0 +1,110 @@ +/** + * `distilly skill migrate` — bring v3 Skill directories up to the current schema. + * + * Idempotent and non-destructive: it creates the v4 layout and seeds an empty + * ledger, and edits nothing except the `schema_version` fields. `--dry-run` + * reports exactly what would change without writing. + */ + +import { relative } from "node:path"; + +import { register } from "./index.mjs"; +import { findSkillDirs, migrateSkillDir, readSchemaVersion, TARGET_SCHEMA_VERSION } from "../skill/migrate.mjs"; + +const help = { + zh: [ + "用法 / Usage:", + " distilly skill migrate [--base-dir

] [--dry-run] [--json]", + "", + "把 v3 目录补成当前 schema(v4):建 knowledge/raw、knowledge/text、evidence/derived、", + "evidence/renders、views,并写一个空账本;**不改任何正文**,只改 schema_version 字段。", + "跑两次是幂等的:第二次不会写任何文件。", + ].join("\n"), + en: [ + "Usage:", + " distilly skill migrate [--base-dir ] [--dry-run] [--json]", + "", + "Brings v3 directories up to the current schema: creates knowledge/raw, knowledge/text,", + "evidence/derived, evidence/renders and views, seeds an empty ledger, and touches nothing", + "but the schema_version fields. Running it twice writes nothing the second time.", + ].join("\n"), +}; + +function parseMigrateArgs(argv) { + const options = { baseDir: process.cwd(), dryRun: false, json: false }; + for (let index = 0; index < argv.length; index += 1) { + const arg = argv[index]; + if (arg === "--json") options.json = true; + else if (arg === "--dry-run") options.dryRun = true; + else if (arg === "--base-dir") { + const value = argv[index + 1]; + if (!value) return { error: "--base-dir requires a value" }; + options.baseDir = value; + index += 1; + } else if (arg.startsWith("--")) return { error: `unknown option: ${arg}` }; + } + return { options }; +} + +register("skill migrate", { + summary: "把 v3 目录迁移到当前 schema / migrate Skill directories", + usage: "distilly skill migrate [--base-dir ] [--dry-run] [--json]", + ...help, + run({ argv, json, reporter }) { + const parsed = parseMigrateArgs(argv); + if (parsed.error) { + return { + receipt: { + command: "skill migrate", + person: null, + ok: false, + inputs: [], + outputs: [], + anchors: { total: 0, cited: 0 }, + warnings: [], + unavailable: [], + error: { code: "skill-migrate/usage", message: parsed.error, remedy: "distilly skill migrate --help" }, + }, + exitCode: 2, + }; + } + const { options } = parsed; + const dirs = findSkillDirs(options.baseDir); + const skills = dirs.map((dir) => { + const result = migrateSkillDir(dir, { dryRun: options.dryRun }); + return { + path: relative(process.cwd(), dir) || ".", + from: result.from, + to: result.to, + changed: result.changed, + actions: result.actions, + }; + }); + const changed = skills.filter((skill) => skill.changed); + const receipt = { + command: "skill migrate", + person: null, + ok: true, + target_schema_version: TARGET_SCHEMA_VERSION, + dry_run: options.dryRun, + skills, + inputs: [], + outputs: [], + anchors: { total: 0, cited: 0 }, + warnings: dirs.length === 0 ? [`no Skill directory found under ${options.baseDir}`] : [], + unavailable: [], + }; + if (!json) { + reporter.line( + `skill migrate${options.dryRun ? " (dry run)" : ""}: ${changed.length}/${skills.length} director(ies) ${options.dryRun ? "would change" : "changed"}`, + ); + for (const skill of changed) { + reporter.line(` ${skill.path}: ${skill.from} → ${skill.to}`); + for (const action of skill.actions) reporter.line(` · ${action}`); + } + } + return { receipt, exitCode: 0 }; + }, +}); + +export const __internal = { readSchemaVersion }; diff --git a/src/commands/skill.mjs b/src/commands/skill.mjs new file mode 100644 index 00000000..7525a237 --- /dev/null +++ b/src/commands/skill.mjs @@ -0,0 +1,480 @@ +/** + * `distilly skill ...` — create / update / list / archive generated Skills. + * + * This module owns the argument surface and the *user-visible* behaviour; the + * work lives in `src/skill/*`, the direct port of `tools/skill_presets.py`, + * `tools/skill_schema.py`, `tools/skill_writer.py` and `tools/version_manager.py`. + * + * Human output is byte-identical to the Python CLI (verified by + * `scripts/parity.mjs`, phase B); `--json` answers with the CONTRACT §3 receipt + * instead, so machine output never has to parse prose. + */ + +import { existsSync, readFileSync, statSync } from "node:fs"; +import { homedir } from "node:os"; +import { join } from "node:path"; + +import { register } from "./index.mjs"; +import { CliError, createReceipt, describeFile, displayPath } from "../cli/receipt.mjs"; +import { parseArgs } from "../cli/args.mjs"; +import { getCharacterPreset, normalizeCharacter, normalizeResearchProfile, resolveExistingStorageRoot } from "../skill/presets.mjs"; +import { resolveContainedChild, validatePathSegment } from "../skill/schema.mjs"; +import { + createSkill, + installGeneratedHosts, + listSkills, + resolveBaseDir, + slugify, + updateSkill, + validateSlug, +} from "../skill/writer.mjs"; +import { SlugResolutionError } from "../skill/slug.mjs"; +import { + backupCurrentVersion, + cleanupOldVersions, + listVersions, + rollback, + MAX_VERSIONS, +} from "../skill/versions.mjs"; + +export function skillHelp(binary = "distilly") { + const zh = [ + "用法:", + ` ${binary} skill create --character [--slug |--name ]`, + " [--meta ] [--work ] [--persona ]", + " [--base-dir ] [--research-profile ]", + " [--install-claude-skill] [--install-openclaw-skill] [--install-codex-skill]", + ` ${binary} skill update --slug [--character ] [--base-dir ]`, + " [--work-patch ] [--persona-patch ] [--correction-json ]", + ` ${binary} skill list [--character ] [--base-dir ]`, + ` ${binary} skill version --slug [--version ]`, + "", + "说明:", + " create 写出 SKILL.md / work.md / persona.md / work_skill.md / persona_skill.md / manifest.json / meta.json。", + " 不带 --slug 时用 --name 生成拼音 slug(Unihan 表);表缺失或字符未覆盖时明确失败并要求 --slug。", + " update 先把当前产物归档到 versions/<当前版本>/,版本号 +1。", + " version 管理归档:list / backup / rollback / cleanup(默认保留最近 10 个)。", + ].join("\n"); + const en = [ + "Usage:", + ` ${binary} skill create --character [--slug |--name ]`, + " [--meta ] [--work ] [--persona ]", + " [--base-dir ] [--research-profile ]", + " [--install-claude-skill] [--install-openclaw-skill] [--install-codex-skill]", + ` ${binary} skill update --slug [--character ] [--base-dir ]`, + " [--work-patch ] [--persona-patch ] [--correction-json ]", + ` ${binary} skill list [--character ] [--base-dir ]`, + ` ${binary} skill version --slug [--version ]`, + "", + "Notes:", + " create writes SKILL.md / work.md / persona.md / work_skill.md / persona_skill.md / manifest.json / meta.json.", + " Without --slug the slug is derived from --name through the Unihan pinyin table; a missing table or an uncovered character fails loudly and asks for --slug.", + " update archives the current artifacts under versions// first, then bumps the version.", + " version manages the archive: list / backup / rollback / cleanup (keeps the newest 10 by default).", + ].join("\n"); + return { zh, en }; +} + +const INSTALL_OPTIONS = { + "install-claude-skill": { type: "boolean" }, + "no-install-claude-skill": { type: "boolean" }, + "install-claude-command-shim": { type: "boolean" }, + "claude-skills-dir": { type: "string", value: "dir" }, + "claude-commands-dir": { type: "string", value: "dir" }, + "install-openclaw-skill": { type: "boolean" }, + "openclaw-skills-dir": { type: "string", value: "dir" }, + "install-codex-skill": { type: "boolean" }, + "codex-skills-dir": { type: "string", value: "dir" }, +}; + +const SELECT_OPTIONS = { + character: { type: "string", alias: "c", value: "family" }, + type: { type: "string", value: "family" }, + "base-dir": { type: "string", value: "dir" }, +}; + +const CREATE_OPTIONS = { + ...SELECT_OPTIONS, + ...INSTALL_OPTIONS, + slug: { type: "string", value: "slug" }, + name: { type: "string", value: "name" }, + meta: { type: "string", value: "file" }, + work: { type: "string", value: "file" }, + persona: { type: "string", value: "file" }, + "research-profile": { type: "string", value: "name" }, +}; + +const UPDATE_OPTIONS = { + ...SELECT_OPTIONS, + ...INSTALL_OPTIONS, + slug: { type: "string", value: "slug" }, + "work-patch": { type: "string", value: "file" }, + "persona-patch": { type: "string", value: "file" }, + "correction-json": { type: "string", value: "file" }, +}; + +const VERSION_OPTIONS = { + ...SELECT_OPTIONS, + slug: { type: "string", value: "slug" }, + version: { type: "string", value: "vN" }, + "max-versions": { type: "string", value: "N" }, +}; + +function readTextFlag(value, label) { + try { + return readFileSync(value, "utf8"); + } catch (error) { + throw new CliError(`cannot read ${label}: ${value}`, { + code: "missing-input", + remedy: error.message, + }); + } +} + +function metaInputs(paths) { + return paths.filter(Boolean).map((path) => describeFile(path)).filter(Boolean); +} + +function artifactOutputs(skillDir) { + const names = [ + "SKILL.md", + "work.md", + "persona.md", + "work_skill.md", + "persona_skill.md", + "manifest.json", + "meta.json", + ]; + return names + .map((name) => describeFile(join(skillDir, name))) + .filter(Boolean); +} + +function resolveSkillDir(baseDir, slug, label = "skill slug") { + let skillDir; + try { + skillDir = resolveContainedChild(baseDir, slug, label); + } catch (error) { + throw new CliError(error.message, { + code: "unsafe-path", + remedy: "pass a single safe directory name (no separators, no '..').", + }); + } + if (!existsSync(skillDir)) { + throw new CliError(`skill directory not found: ${skillDir}`, { + code: "missing-skill", + remedy: "run `distilly skill list` to see the skills in this storage root.", + }); + } + return skillDir; +} + +const createCommand = { + summary: "创建 Skill / Create a Skill", + usage: "distilly skill create [options]", + options: CREATE_OPTIONS, + ...skillHelp(), + run({ argv, json, reporter }) { + const { flags } = parseArgs(argv, CREATE_OPTIONS); + const requestedCharacter = normalizeCharacter(flags.character || flags.type); + + const autoInstallSetting = + process.env.DISTILLY_AUTO_INSTALL_CLAUDE ?? process.env.DOT_SKILL_AUTO_INSTALL_CLAUDE; + const autoInstallDefault = autoInstallSetting !== undefined && autoInstallSetting !== "0"; + const installClaudeSkill = + (flags["install-claude-skill"] || autoInstallDefault) && !flags["no-install-claude-skill"]; + + const meta = flags.meta ? JSON.parse(readTextFlag(flags.meta, "meta JSON")) : {}; + if (flags.name) { + meta.name = flags.name; + meta.display_name = flags.name; + } + meta.character = normalizeCharacter(meta.character ?? meta.type ?? requestedCharacter); + meta.research_profile = normalizeResearchProfile( + meta.character, + flags["research-profile"] || meta.research_profile, + ); + meta.type = meta.type || meta.character; + + const baseDir = resolveBaseDir(flags["base-dir"], requestedCharacter); + let slug; + try { + slug = flags.slug + ? validateSlug(flags.slug) + : slugify(meta.display_name ?? meta.name ?? "person"); + } catch (error) { + if (error instanceof SlugResolutionError) { + throw new CliError(error.message, { code: error.code, remedy: error.remedy }); + } + throw new CliError(error.message, { + code: "invalid-slug", + remedy: "pass --slug (1-40 lowercase letters/digits).", + }); + } + + const workContent = flags.work ? readTextFlag(flags.work, "work.md") : ""; + const personaContent = flags.persona ? readTextFlag(flags.persona, "persona.md") : ""; + + const skillDir = createSkill(baseDir, slug, meta, workContent, personaContent); + + reporter.line(`Created skill: ${displayPath(skillDir)}`); + reporter.line(" Kind: meta-skill"); + reporter.line(` Character: ${meta.character}`); + reporter.line(` Research Profile: ${meta.research_profile}`); + reporter.line(` Preset: ${meta.preset ?? "auto"}`); + + const installLines = installGeneratedHosts( + skillDir, + { + claudeSkillsDir: flags["claude-skills-dir"], + claudeCommandsDir: flags["claude-commands-dir"], + installClaudeCommandShim: flags["install-claude-command-shim"], + openclawSkillsDir: flags["openclaw-skills-dir"], + installOpenclawSkill: flags["install-openclaw-skill"], + codexSkillsDir: flags["codex-skills-dir"], + installCodexSkill: flags["install-codex-skill"], + }, + installClaudeSkill, + ); + if (installLines.length > 0) for (const line of installLines) reporter.line(line); + else reporter.line(" Host installs: skipped"); + + const outputs = artifactOutputs(skillDir); + const warnings = []; + if (!json && flags.slug === undefined && slug) { + warnings.push(`slug derived from --name: ${slug}`); + } + return { + receipt: createReceipt("skill create", { + person: slug, + inputs: metaInputs([flags.meta, flags.work, flags.persona]), + outputs, + warnings, + }), + }; + }, +}; + +const updateCommand = { + summary: "更新 Skill / Update a Skill", + usage: "distilly skill update [options]", + options: UPDATE_OPTIONS, + ...skillHelp(), + run({ argv, reporter }) { + const { flags } = parseArgs(argv, UPDATE_OPTIONS); + const requestedCharacter = normalizeCharacter(flags.character || flags.type); + + let slug; + try { + slug = validatePathSegment(flags.slug ?? "", "existing slug"); + } catch (error) { + throw new CliError(error.message, { + code: "unsafe-slug", + remedy: "pass --slug .", + }); + } + + const baseDir = resolveExistingStorageRoot(requestedCharacter, slug, flags["base-dir"]); + const skillDir = resolveSkillDir(baseDir, slug); + + const workPatch = flags["work-patch"] ? readTextFlag(flags["work-patch"], "work patch") : null; + const personaPatch = flags["persona-patch"] + ? readTextFlag(flags["persona-patch"], "persona patch") + : null; + const correction = flags["correction-json"] + ? JSON.parse(readTextFlag(flags["correction-json"], "correction JSON")) + : null; + + const newVersion = updateSkill(skillDir, workPatch, personaPatch, correction); + reporter.line(`Updated skill to ${newVersion}: ${displayPath(skillDir)}`); + + const autoInstallSetting = + process.env.DISTILLY_AUTO_INSTALL_CLAUDE ?? process.env.DOT_SKILL_AUTO_INSTALL_CLAUDE; + const autoInstallDefault = autoInstallSetting !== undefined && autoInstallSetting !== "0"; + const installClaudeSkill = + (flags["install-claude-skill"] || autoInstallDefault) && !flags["no-install-claude-skill"]; + const installLines = installGeneratedHosts( + skillDir, + { + claudeSkillsDir: flags["claude-skills-dir"], + claudeCommandsDir: flags["claude-commands-dir"], + installClaudeCommandShim: flags["install-claude-command-shim"], + openclawSkillsDir: flags["openclaw-skills-dir"], + installOpenclawSkill: flags["install-openclaw-skill"], + codexSkillsDir: flags["codex-skills-dir"], + installCodexSkill: flags["install-codex-skill"], + }, + installClaudeSkill, + ); + for (const line of installLines) reporter.line(line); + + return { + receipt: createReceipt("skill update", { + person: slug, + inputs: metaInputs([ + flags["work-patch"], + flags["persona-patch"], + flags["correction-json"], + join(skillDir, "meta.json"), + ]), + outputs: artifactOutputs(skillDir), + warnings: [], + }), + }; + }, +}; + +const listCommand = { + summary: "列出已有 Skill / List generated Skills", + usage: "distilly skill list [options]", + options: SELECT_OPTIONS, + ...skillHelp(), + run({ argv, reporter }) { + const { flags } = parseArgs(argv, SELECT_OPTIONS); + const requestedCharacter = normalizeCharacter(flags.character || flags.type); + const baseDir = resolveExistingStorageRoot(requestedCharacter, null, flags["base-dir"]); + const skills = listSkills(baseDir); + + if (skills.length === 0) { + const preset = getCharacterPreset(requestedCharacter); + reporter.line(`No ${preset.character} skills found`); + } else { + reporter.line(`Found ${skills.length} skills:`); + reporter.line(""); + for (const skill of skills) { + const updated = skill.updated_at ? skill.updated_at.slice(0, 10) : "unknown"; + reporter.line(` [${skill.slug}] ${skill.name} — ${skill.identity}`); + reporter.line( + ` Kind: ${skill.kind} Character: ${skill.character} ` + + `Research Profile: ${skill.research_profile} ` + + `Version: ${skill.version} ` + + `Corrections: ${skill.corrections_count} Updated: ${updated}`, + ); + reporter.line(""); + } + } + + const outputs = skills + .map((skill) => describeFile(join(baseDir, skill.slug, "SKILL.md"))) + .filter(Boolean); + return { + receipt: createReceipt("skill list", { + outputs, + warnings: skills.length === 0 ? [`no skills found in ${baseDir}`] : [], + }), + }; + }, +}; + +const versionCommand = { + summary: "版本归档 / Archive, roll back and prune Skill versions", + usage: "distilly skill version [options]", + options: VERSION_OPTIONS, + ...skillHelp(), + run({ argv, reporter }) { + const { flags, positionals } = parseArgs(argv, VERSION_OPTIONS); + const action = positionals[0] ?? "list"; + if (!["list", "backup", "rollback", "cleanup"].includes(action)) { + throw new CliError(`unknown skill version action: ${action}`, { + code: "usage", + remedy: "choose one of: list, backup, rollback, cleanup.", + }); + } + + const requestedCharacter = normalizeCharacter(flags.character || flags.type); + let slug; + try { + slug = validatePathSegment(flags.slug ?? "", "skill slug"); + } catch (error) { + throw new CliError(error.message, { + code: "unsafe-slug", + remedy: "pass --slug .", + }); + } + const baseDir = resolveExistingStorageRoot(requestedCharacter, slug, flags["base-dir"]); + const skillDir = resolveSkillDir(baseDir, slug); + + let outputs = []; + if (action === "list") { + const versions = listVersions(skillDir); + if (versions.length === 0) { + reporter.line(`no archived versions for ${slug}`); + } else { + reporter.line(`archived versions for ${slug}:`); + reporter.line(""); + for (const version of versions) { + reporter.line( + ` ${version.version} archived: ${version.archived_at} files: ${version.files.join(", ")}`, + ); + } + } + outputs = versions + .map((version) => describeFile(join(version.path, "SKILL.md"))) + .filter(Boolean); + } else if (action === "backup") { + if (!backupCurrentVersion(skillDir)) { + throw new CliError(`could not archive the current version of ${slug}`, { + code: "archive-failed", + remedy: "make sure meta.json exists in the skill directory.", + }); + } + outputs = artifactOutputs(skillDir); + } else if (action === "rollback") { + if (!flags.version) { + throw new CliError("rollback requires --version", { + code: "usage", + remedy: "run `distilly skill version list --slug ` and pass --version .", + }); + } + if (!rollback(skillDir, flags.version)) { + throw new CliError(`rollback to ${flags.version} failed`, { + code: "rollback-failed", + remedy: `check the archive list for ${slug}.`, + }); + } + outputs = artifactOutputs(skillDir); + } else { + const maxVersions = flags["max-versions"] ? Number.parseInt(flags["max-versions"], 10) : MAX_VERSIONS; + if (!cleanupOldVersions(skillDir, maxVersions)) { + throw new CliError(`cleanup failed for ${slug}`, { code: "cleanup-failed" }); + } + reporter.line("cleanup complete"); + outputs = []; + } + + return { + receipt: createReceipt("skill version", { + person: slug, + inputs: [describeFile(join(skillDir, "meta.json"))].filter(Boolean), + outputs, + warnings: [], + }), + }; + }, +}; + +export function registerSkillCommands() { + register("skill", { + summary: "Skill 子命令入口 / Skill subcommand entry", + usage: "distilly skill [options]", + ...skillHelp(), + run({ argv, reporter }) { + if (argv.length === 0) { + reporter.line(skillHelp().zh); + return { receipt: createReceipt("skill", { warnings: [] }) }; + } + throw new CliError(`unknown skill subcommand: ${argv[0]}`, { + code: "usage", + remedy: "choose one of: create, update, list, version.", + }); + }, + }); + register("skill create", createCommand); + register("skill update", updateCommand); + register("skill list", listCommand); + register("skill version", versionCommand); +} + +registerSkillCommands(); + +export { homedir, statSync }; diff --git a/src/skill/migrate.mjs b/src/skill/migrate.mjs new file mode 100644 index 00000000..bd8848a7 --- /dev/null +++ b/src/skill/migrate.mjs @@ -0,0 +1,118 @@ +/** + * Schema migrations for generated Skill directories. + * + * v4 adds the v2 evidence layout — `knowledge/raw`, `knowledge/text`, + * `knowledge/index.json`, `evidence/derived`, `evidence/renders`, `views` — to + * directories created by the v3 engine. It never rewrites a body: the six + * primary artifacts and every markdown file keep their bytes, and the only edits + * are the `schema_version` fields plus the directories and an empty ledger. + * + * Migration is idempotent by construction (a directory that already reports v4 is + * returned untouched), so running it twice is a no-op the second time. + */ + +import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs"; +import { join } from "node:path"; + +import { SCHEMA_VERSION, jsonDumps } from "./schema.mjs"; +import { KnowledgeStore } from "../knowledge/store.mjs"; +import { emptyLedger, saveLedger } from "../knowledge/ledger.mjs"; + +export const TARGET_SCHEMA_VERSION = SCHEMA_VERSION; + +/** Directories v4 guarantees; the ledger is seeded separately. */ +export const V4_DIRECTORIES = [ + "knowledge/raw", + "knowledge/text", + "evidence/derived", + "evidence/renders", + "views", +]; + +function readJson(path) { + return JSON.parse(readFileSync(path, "utf8")); +} + +/** `schema_version` as reported by meta.json, then manifest.json, else "3". */ +export function readSchemaVersion(skillDir) { + for (const file of ["meta.json", "manifest.json"]) { + const path = join(skillDir, file); + if (!existsSync(path)) continue; + try { + const value = readJson(path).schema_version; + if (typeof value === "string" && value.length > 0) return value; + } catch { + /* a malformed artifact is reported by `doctor`, not by the migration */ + } + } + return "3"; +} + +/** Rewrite only `schema_version`, preserving key order and the write format. */ +function bumpArtifactVersion(path, version, actions, dryRun, relative) { + if (!existsSync(path)) return; + const value = readJson(path); + if (value.schema_version === version) return; + value.schema_version = version; + if (!dryRun) writeFileSync(path, jsonDumps(value), "utf8"); + actions.push(`bumped schema_version in ${relative}`); +} + +/** + * Migrate one Skill directory to the current schema. + * @returns {{skillDir: string, from: string, to: string, actions: string[], changed: boolean}} + */ +export function migrateSkillDir(skillDir, { dryRun = false } = {}) { + const from = readSchemaVersion(skillDir); + const actions = []; + if (from === TARGET_SCHEMA_VERSION) { + return { skillDir, from, to: from, actions, changed: false }; + } + + for (const relative of V4_DIRECTORIES) { + const target = join(skillDir, relative); + if (existsSync(target)) continue; + if (!dryRun) mkdirSync(target, { recursive: true }); + actions.push(`created ${relative}`); + } + + const store = new KnowledgeStore(skillDir, { dryRun }); + if (!existsSync(store.ledgerPath)) { + if (!dryRun) saveLedger(store, emptyLedger()); + actions.push("seeded knowledge/index.json"); + } + + bumpArtifactVersion(join(skillDir, "meta.json"), TARGET_SCHEMA_VERSION, actions, dryRun, "meta.json"); + bumpArtifactVersion(join(skillDir, "manifest.json"), TARGET_SCHEMA_VERSION, actions, dryRun, "manifest.json"); + + return { skillDir, from, to: TARGET_SCHEMA_VERSION, actions, changed: actions.length > 0 }; +} + +/** Every `skills//` directory under `baseDir` (depth ≤ 2). */ +export function findSkillDirs(baseDir) { + const found = []; + const familyRoot = join(baseDir, "skills"); + const root = existsSync(familyRoot) ? familyRoot : baseDir; + for (const family of safeReaddir(root)) { + const familyPath = join(root, family); + for (const slug of safeReaddir(familyPath)) { + const candidate = join(familyPath, slug); + if (existsSync(join(candidate, "meta.json")) || existsSync(join(candidate, "manifest.json"))) { + found.push(candidate); + } + } + } + return found.sort(); +} + +function safeReaddir(dir) { + try { + return require("node:fs") + .readdirSync(dir, { withFileTypes: true }) + .filter((entry) => entry.isDirectory() && !entry.name.startsWith(".")) + .map((entry) => entry.name) + .sort(); + } catch { + return []; + } +} diff --git a/tests/cli-lifecycle.test.mjs b/tests/cli-lifecycle.test.mjs new file mode 100644 index 00000000..6ede3128 --- /dev/null +++ b/tests/cli-lifecycle.test.mjs @@ -0,0 +1,464 @@ +/** + * Port of `tests/test_cli_lifecycle.py` driven through `bin/distilly.mjs`. + * + * The Python suite shelled out to `python3 tools/skill_writer.py` and + * `python3 tools/version_manager.py`; this port runs the equivalent Node CLI + * commands. The celebrity branch's research-tool steps + * (`tools/research/srt_to_transcript.py`, `merge_research.py`, `quality_check.py`) + * stay in Python until ds/02-parse-zero-cred ports them, and are therefore not + * exercised here — see docs/evidence/pr-01-node-core.md. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, renameSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); +const cli = join(projectRoot, "bin", "distilly.mjs"); + +function runCmd(args, { cwd = projectRoot, env = {} } = {}) { + const merged = { ...process.env, DISTILLY_AUTO_INSTALL_CLAUDE: "0", ...env }; + return spawnSync(process.execPath, [cli, ...args], { cwd, encoding: "utf8", env: merged }); +} + +function writeJson(path, payload) { + writeFileSync(path, JSON.stringify(payload, null, 2), "utf8"); + return path; +} + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-cli-")); +} + +test("the default colleague CLI uses the skills/colleague root", () => { + const root = tempDir(); + try { + const workPath = join(root, "work.md"); + const personaPath = join(root, "persona.md"); + writeFileSync(workPath, "Work body\n", "utf8"); + writeFileSync(personaPath, "Persona body\n", "utf8"); + const metaPath = writeJson(join(root, "meta.json"), { + character: "colleague", + display_name: "Eulalie", + classification: { language: "en" }, + }); + + const create = runCmd( + [ + "skill", + "create", + "--character", + "colleague", + "--slug", + "eulalie", + "--name", + "Eulalie", + "--meta", + metaPath, + "--work", + workPath, + "--persona", + personaPath, + ], + { cwd: root }, + ); + + assert.equal(create.status, 0, create.stderr); + assert.match(create.stdout, /Created skill:/); + assert.ok(existsSync(join(root, "skills", "colleague", "eulalie", "SKILL.md"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("Claude auto-install is opt-in and keeps the legacy env variable working", () => { + const root = tempDir(); + try { + const home = join(root, "home"); + const baseDir = join(root, "skills", "colleague"); + mkdirSync(home); + mkdirSync(baseDir, { recursive: true }); + const metaPath = writeJson(join(root, "meta.json"), { + character: "colleague", + display_name: "Eulalie", + classification: { language: "en" }, + }); + const workPath = join(root, "work.md"); + const personaPath = join(root, "persona.md"); + writeFileSync(workPath, "Work body\n", "utf8"); + writeFileSync(personaPath, "Persona body\n", "utf8"); + + const createWithEnv = (slug, settings, ...extraArgs) => { + const env = { ...process.env, HOME: home }; + delete env.DISTILLY_AUTO_INSTALL_CLAUDE; + delete env.DOT_SKILL_AUTO_INSTALL_CLAUDE; + Object.assign(env, settings); + const result = spawnSync( + process.execPath, + [ + cli, + "skill", + "create", + "--character", + "colleague", + "--slug", + slug, + "--name", + "Eulalie", + "--meta", + metaPath, + "--work", + workPath, + "--persona", + personaPath, + "--base-dir", + baseDir, + ...extraArgs, + ], + { cwd: projectRoot, encoding: "utf8", env }, + ); + assert.equal(result.status, 0, result.stderr); + return join(home, ".claude", "skills", `colleague-${slug}`, "SKILL.md"); + }; + + assert.equal(existsSync(createWithEnv("default-off", {})), false); + assert.equal(existsSync(createWithEnv("legacy-on", { DOT_SKILL_AUTO_INSTALL_CLAUDE: "1" })), true); + assert.equal(existsSync(createWithEnv("legacy-off", { DOT_SKILL_AUTO_INSTALL_CLAUDE: "0" })), false); + assert.equal( + existsSync( + createWithEnv("new-wins-off", { + DISTILLY_AUTO_INSTALL_CLAUDE: "0", + DOT_SKILL_AUTO_INSTALL_CLAUDE: "1", + }), + ), + false, + ); + assert.equal( + existsSync( + createWithEnv("new-wins-on", { + DISTILLY_AUTO_INSTALL_CLAUDE: "1", + DOT_SKILL_AUTO_INSTALL_CLAUDE: "0", + }), + ), + true, + ); + assert.equal( + existsSync(createWithEnv("explicit-off", { DISTILLY_AUTO_INSTALL_CLAUDE: "1" }, "--no-install-claude-skill")), + false, + ); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create with only --name normalises the slug and rejects an unsafe explicit slug", () => { + const root = tempDir(); + try { + const created = runCmd(["skill", "create", "--name", "Zadie Smith", "--base-dir", "skills/colleague"], { + cwd: root, + }); + assert.equal(created.status, 0, created.stderr); + const generated = join(root, "skills", "colleague", "zadie-smith", "SKILL.md"); + assert.match(readFileSync(generated, "utf8"), /name: colleague-zadie-smith/); + + const unsafe = runCmd(["skill", "create", "--slug", "../escape", "--base-dir", "skills/colleague"], { + cwd: root, + }); + assert.notEqual(unsafe.status, 0); + assert.equal(existsSync(join(root, "skills", "escape")), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update accepts a safe legacy slug with spaces", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const create = runCmd( + ["skill", "create", "--slug", "legacy", "--name", "Zadie Smith", "--base-dir", baseDir], + { cwd: root }, + ); + assert.equal(create.status, 0, create.stderr); + const legacyDir = join(baseDir, "Zadie Smith"); + renameSync(join(baseDir, "legacy"), legacyDir); + const workPatch = join(root, "work-patch.md"); + writeFileSync(workPatch, "## Update\n\nLegacy directory remains addressable.\n", "utf8"); + + const update = runCmd( + ["skill", "update", "--slug", "Zadie Smith", "--base-dir", baseDir, "--work-patch", workPatch], + { cwd: root }, + ); + + assert.equal(update.status, 0, update.stderr); + assert.match(update.stdout, /Updated skill to v2:/); + assert.match(update.stdout, /Zadie Smith/); + assert.match(readFileSync(join(legacyDir, "work.md"), "utf8"), /Legacy directory remains addressable/); + const savedMeta = JSON.parse(readFileSync(join(legacyDir, "meta.json"), "utf8")); + assert.equal(savedMeta.artifacts.combined_command, "colleague-zadie-smith"); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("the version manager rejects slug and version traversal", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const victimVersions = join(root, "skills", "victim", "versions"); + for (let index = 0; index < 11; index += 1) mkdirSync(join(victimVersions, `v${index}`), { recursive: true }); + + const traversal = runCmd(["skill", "version", "cleanup", "--slug", "../victim", "--base-dir", baseDir], { + cwd: root, + }); + assert.notEqual(traversal.status, 0); + assert.equal(readFileSync(join(victimVersions, "..", "..", "skills", "victim", "versions")) ? 11 : 0, 11); + + const create = runCmd(["skill", "create", "--slug", "safe", "--name", "Safe", "--base-dir", baseDir], { + cwd: root, + }); + assert.equal(create.status, 0, create.stderr); + const rollbackTraversal = runCmd( + ["skill", "version", "rollback", "--slug", "safe", "--version", "../victim", "--base-dir", baseDir], + { cwd: root }, + ); + assert.notEqual(rollbackTraversal.status, 0); + assert.ok(existsSync(join(baseDir, "safe", "SKILL.md"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("each character family survives the full CLI lifecycle", () => { + const fixtures = { + colleague: { name: "Eulalie", slug: "eulalie", baseDir: "skills/colleague" }, + relationship: { name: "Mireille", slug: "mireille", baseDir: "skills/relationship" }, + celebrity: { name: "Zadie Smith", slug: "zadie-smith", baseDir: "skills/celebrity" }, + }; + + const root = tempDir(); + try { + for (const [character, fixture] of Object.entries(fixtures)) { + const baseDir = join(root, fixture.baseDir); + mkdirSync(baseDir, { recursive: true }); + const metaPath = writeJson(join(root, `${fixture.slug}_meta.json`), { + character, + display_name: fixture.name, + classification: { language: "en" }, + profile: { role: "Builder" }, + tags: { personality: ["precise", "skeptical"] }, + knowledge_sources: ["manual-notes"], + }); + const workPath = join(root, `${fixture.slug}_work.md`); + const personaPath = join(root, `${fixture.slug}_persona.md`); + const workPatchPath = join(root, `${fixture.slug}_work_patch.md`); + const correctionPath = join(root, `${fixture.slug}_correction.json`); + + writeFileSync( + workPath, + [ + "## mental models", + "- First-principles reasoning", + "- Skeptical framing", + "- Long-horizon tradeoffs", + "", + "## limitations", + "- Avoids operational detail", + "", + "Sources:", + "https://example.com/articles/long-form-profile", + "https://example.com/interviews/episode-42", + ].join("\n") + "\n", + "utf8", + ); + writeFileSync( + personaPath, + [ + "## expression DNA", + "- Sentence rhythm is clipped.", + "- Uses metaphor when disagreeing.", + "", + "## honest boundaries", + "- States what they do not know.", + "", + "## contradictions", + "- Alternates between certainty and doubt.", + ].join("\n") + "\n", + "utf8", + ); + writeFileSync(workPatchPath, "## new evidence\n- Adds a later example.\n", "utf8"); + writeJson(correctionPath, { + scene: "disagreement", + wrong: "flatten disagreement into politeness", + correct: "surface the disagreement and justify it directly", + }); + + const create = runCmd( + [ + "skill", + "create", + "--character", + character, + "--slug", + fixture.slug, + "--name", + fixture.name, + "--meta", + metaPath, + "--work", + workPath, + "--persona", + personaPath, + "--base-dir", + baseDir, + ], + { cwd: root }, + ); + assert.equal(create.status, 0, create.stderr); + assert.match(create.stdout, /Created skill:/); + + const skillDir = join(baseDir, fixture.slug); + assert.ok(existsSync(join(skillDir, "SKILL.md"))); + assert.ok(existsSync(join(skillDir, "manifest.json"))); + + const listResult = runCmd(["skill", "list", "--character", character, "--base-dir", baseDir], { cwd: root }); + assert.match(listResult.stdout, new RegExp(fixture.slug)); + assert.match(listResult.stdout, new RegExp(`Character: ${character}`)); + + const update = runCmd( + [ + "skill", + "update", + "--character", + character, + "--slug", + fixture.slug, + "--work-patch", + workPatchPath, + "--correction-json", + correctionPath, + "--base-dir", + baseDir, + ], + { cwd: root }, + ); + assert.equal(update.status, 0, update.stderr); + assert.match(update.stdout, /Updated skill to v2/); + + const versions = runCmd( + ["skill", "version", "list", "--character", character, "--slug", fixture.slug, "--base-dir", baseDir], + { cwd: root }, + ); + assert.match(versions.stdout, /v1/); + + const rollbackResult = runCmd( + [ + "skill", + "version", + "rollback", + "--character", + character, + "--slug", + fixture.slug, + "--version", + "v1", + "--base-dir", + baseDir, + ], + { cwd: root }, + ); + assert.equal(rollbackResult.status, 0, rollbackResult.stderr); + assert.match(rollbackResult.stdout, /rolled back to v1/); + + const savedMeta = JSON.parse(readFileSync(join(skillDir, "meta.json"), "utf8")); + assert.equal(savedMeta.character, character); + assert.ok(savedMeta.version.startsWith("v1")); + + const combinedSkill = readFileSync(join(skillDir, "SKILL.md"), "utf8"); + assert.match(combinedSkill, /## PART A: Work/); + assert.match(combinedSkill, /## PART B: Persona/); + + if (character === "celebrity") { + // The research toolchain (srt_to_transcript / merge_research / + // quality_check) still ships as Python; the directory layout is what + // this branch owns, so assert that instead. + assert.ok(existsSync(join(skillDir, "knowledge", "subtitles"))); + assert.ok(existsSync(join(skillDir, "knowledge", "transcripts"))); + } + } + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("the CLI installs a generated skill into the supported host paths", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + mkdirSync(baseDir, { recursive: true }); + + const metaPath = writeJson(join(root, "zhou_qimo_meta.json"), { + character: "celebrity", + display_name: "周奇墨", + classification: { language: "zh-CN" }, + }); + const workPath = join(root, "zhou_qimo_work.md"); + const personaPath = join(root, "zhou_qimo_persona.md"); + writeFileSync(workPath, "Work body\n", "utf8"); + writeFileSync(personaPath, "Persona body\n", "utf8"); + + const claudeSkillsDir = join(root, ".claude", "skills"); + const claudeCommandsDir = join(root, ".claude", "commands"); + const openclawSkillsDir = join(root, ".openclaw", "workspace", "skills"); + const codexSkillsDir = join(root, ".agents", "skills"); + + const create = runCmd( + [ + "skill", + "create", + "--character", + "celebrity", + "--slug", + "zhou-qimo", + "--name", + "周奇墨", + "--meta", + metaPath, + "--work", + workPath, + "--persona", + personaPath, + "--base-dir", + baseDir, + "--install-claude-skill", + "--install-claude-command-shim", + "--claude-skills-dir", + claudeSkillsDir, + "--claude-commands-dir", + claudeCommandsDir, + "--install-openclaw-skill", + "--openclaw-skills-dir", + openclawSkillsDir, + "--install-codex-skill", + "--codex-skills-dir", + codexSkillsDir, + ], + { cwd: root }, + ); + + assert.equal(create.status, 0, create.stderr); + assert.match(create.stdout, /Claude trigger: \/celebrity-zhou-qimo/); + assert.match(create.stdout, /OpenClaw trigger: \/celebrity-zhou-qimo/); + assert.match(create.stdout, /Codex skill name: celebrity-zhou-qimo/); + assert.ok(existsSync(join(claudeSkillsDir, "celebrity-zhou-qimo", "SKILL.md"))); + assert.ok(existsSync(join(claudeCommandsDir, "celebrity-zhou-qimo.md"))); + assert.ok(existsSync(join(openclawSkillsDir, "celebrity-zhou-qimo", "SKILL.md"))); + assert.ok(existsSync(join(codexSkillsDir, "celebrity-zhou-qimo", "SKILL.md"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); diff --git a/tests/commands.test.mjs b/tests/commands.test.mjs new file mode 100644 index 00000000..a909c6d5 --- /dev/null +++ b/tests/commands.test.mjs @@ -0,0 +1,235 @@ +/** + * Coverage for the CLI surface this branch adds: install / uninstall / doctor / + * legacy forwarding, plus the receipt discipline (`--json` stdout is one object). + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +import { runCli } from "./dispatcher.test.mjs"; +import { supportedHosts } from "../src/install/hosts.mjs"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-cmd-")); +} + +function parseReceipt(stdout) { + const start = stdout.indexOf("{"); + assert.notEqual(start, -1, `no receipt on stdout: ${stdout}`); + return JSON.parse(stdout.slice(start)); +} + +test("install --path copies the payload and reports it with a receipt", () => { + const root = tempDir(); + try { + const target = join(root, "host", "skills", "distilly"); + const result = runCli(["install", "--path", target, "--json"]); + + assert.equal(result.status, 0, result.stderr); + const receipt = parseReceipt(result.stdout); + assert.equal(receipt.command, "install"); + assert.equal(receipt.ok, true); + assert.equal(receipt.host, null); + assert.ok(receipt.outputs.length > 0); + for (const output of receipt.outputs) { + assert.equal(typeof output.path, "string"); + assert.equal(typeof output.sha256, "string"); + assert.equal(typeof output.bytes, "number"); + } + assert.ok(existsSync(join(target, "SKILL.md"))); + assert.ok(existsSync(join(target, "src", "commands", "index.mjs"))); + + // A second install without --force refuses and points at the remedy. + const again = runCli(["install", "--path", target, "--json"]); + assert.notEqual(again.status, 0); + const failed = parseReceipt(again.stdout); + assert.equal(failed.ok, false); + assert.match(failed.error.remedy, /--force/); + + // --force replaces it and preserves the previous copy. + const forced = runCli(["install", "--path", target, "--force"]); + assert.equal(forced.status, 0, forced.stderr); + assert.match(forced.stdout, /Previous install preserved at/); + assert.ok(existsSync(target)); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install refuses a target whose final segment is not distilly", () => { + const root = tempDir(); + try { + const result = runCli(["install", "--path", join(root, "somewhere-else")]); + assert.notEqual(result.status, 0); + assert.match(result.stderr, /must end with a directory named distilly/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install resolves the documented directory from the shared matrix", () => { + const home = tempDir(); + try { + for (const host of supportedHosts()) { + const result = runCli(["install", host, "--dry-run", "--json"], { env: { HOME: home } }); + assert.equal(result.status, 0, `${host}: ${result.stderr}`); + const receipt = parseReceipt(result.stdout); + assert.equal(receipt.host, host); + assert.equal(receipt.dry_run, true); + assert.equal(receipt.outputs.length, 0, "dry run must not write"); + } + // the project scope is only offered where the host documents one + const openclawProject = runCli(["install", "openclaw", "--project", "--dry-run"], { env: { HOME: home } }); + assert.notEqual(openclawProject.status, 0); + assert.match(openclawProject.stderr, /no documented project-local directory/); + } finally { + rmSync(home, { recursive: true, force: true }); + } +}); + +test("uninstall removes an install, keeps a backup on request and refuses strangers", () => { + const root = tempDir(); + try { + const target = join(root, "host", "skills", "distilly"); + assert.equal(runCli(["install", "--path", target]).status, 0); + + const dryRun = runCli(["uninstall", "--path", target, "--dry-run"]); + assert.equal(dryRun.status, 0, dryRun.stderr); + assert.ok(existsSync(target), "dry run must not delete"); + + const removed = runCli(["uninstall", "--path", target, "--backup"]); + assert.equal(removed.status, 0, removed.stderr); + assert.match(removed.stdout, /kept at/); + assert.equal(existsSync(target), false); + + // A directory that is not a Distilly install needs --force. + mkdirSync(target, { recursive: true }); + writeFileSync(join(target, "SKILL.md"), "---\nname: something-else\n---\n", "utf8"); + const refused = runCli(["uninstall", "--path", target]); + assert.notEqual(refused.status, 0); + assert.match(refused.stderr, /not a Distilly install/); + assert.equal(existsSync(target), true); + + const forced = runCli(["uninstall", "--path", target, "--force"]); + assert.equal(forced.status, 0, forced.stderr); + assert.equal(existsSync(target), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("doctor inventories every host and names what this build cannot do yet", () => { + const home = tempDir(); + try { + const result = runCli(["doctor", "--json"], { env: { HOME: home } }); + assert.equal(result.status, 0, result.stderr); + const receipt = parseReceipt(result.stdout); + + assert.equal(receipt.command, "doctor"); + assert.ok(Array.isArray(receipt.hosts)); + assert.deepEqual( + receipt.hosts.map((row) => row.host), + supportedHosts(), + ); + for (const row of receipt.hosts) { + assert.equal(row.installed, false); + assert.equal(typeof row.path, "string"); + } + const unavailable = receipt.unavailable.map((item) => item.channel); + for (const command of ["harvest", "retrospect", "view", "collect"]) { + assert.ok(unavailable.includes(command), `${command} must be reported as unavailable`); + } + assert.deepEqual(receipt.anchors, { total: 0, cited: 0 }); + } finally { + rmSync(home, { recursive: true, force: true }); + } +}); + +test("the legacy adapter forwards old command lines with a deprecation warning", () => { + const root = tempDir(); + try { + writeFileSync(join(root, "work.md"), "Work body\n", "utf8"); + const create = runCli( + [ + "legacy", + "tools/skill_writer.py", + "--action", + "create", + "--character", + "colleague", + "--slug", + "eulalie", + "--name", + "Eulalie", + "--work", + "work.md", + "--base-dir", + "skills/colleague", + ], + { cwd: root }, + ); + assert.equal(create.status, 0, create.stderr); + assert.match(create.stdout, /Created skill:/); + assert.match(create.stderr, /deprecated/); + assert.ok(existsSync(join(root, "skills", "colleague", "eulalie", "SKILL.md"))); + + const list = runCli( + ["legacy", "skill_writer.py", "--action", "list", "--character", "colleague", "--base-dir", "skills/colleague"], + { cwd: root }, + ); + assert.equal(list.status, 0, list.stderr); + assert.match(list.stdout, /Found 1 skills:/); + + const version = runCli( + ["legacy", "version_manager.py", "--action", "backup", "--slug", "eulalie", "--base-dir", "skills/colleague"], + { cwd: root }, + ); + assert.equal(version.status, 0, version.stderr); + assert.match(version.stdout, /archived version v1/); + + const unknown = runCli(["legacy", "feishu_parser.py"]); + assert.notEqual(unknown.status, 0); + assert.match(unknown.stderr, /no legacy adapter/); + + const badAction = runCli(["legacy", "skill_writer.py", "--action", "explode"]); + assert.notEqual(badAction.status, 0); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("--json keeps stdout machine-readable for every command this branch ships", () => { + const root = tempDir(); + try { + const invocations = [ + ["skill", "list", "--character", "colleague", "--base-dir", "skills/colleague", "--json"], + ["doctor", "--json"], + ["install", "codex", "--dry-run", "--json"], + ]; + for (const args of invocations) { + const result = runCli(args, { cwd: root, env: { HOME: root } }); + assert.doesNotThrow(() => JSON.parse(result.stdout), `${args.join(" ")} stdout is not pure JSON`); + assert.equal(result.stdout.trim().startsWith("{"), true); + assert.equal(result.stdout.trim().endsWith("}"), true); + } + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("the packaged payload is complete", () => { + const result = runCli(["--check-package"]); + assert.equal(result.status, 0, result.stderr); + // The verdict goes to stderr: `prepack` shares stdout with `npm pack --json`. + assert.match(result.stderr, /payload is valid/); + assert.equal(result.stdout, "", "`--check-package` must leave stdout clean for machine consumers"); + const skill = readFileSync(join(projectRoot, "SKILL.md"), "utf8"); + const version = JSON.parse(readFileSync(join(projectRoot, "package.json"), "utf8")).version; + assert.ok(skill.includes(`version: "${version}"`)); +}); diff --git a/tests/entry-point.test.mjs b/tests/entry-point.test.mjs new file mode 100644 index 00000000..49a417fe --- /dev/null +++ b/tests/entry-point.test.mjs @@ -0,0 +1,87 @@ +/** + * Entry guards must survive a symlinked path. + * + * The natural spelling — `resolve(process.argv[1]) === fileURLToPath(import.meta.url)` + * — compares a caller-supplied path against a realpath. Under a symlink they + * differ, the file concludes it was imported, and it exits 0 having done nothing. + * Both halves of that are real here: `/tmp` is a symlink to `/private/tmp` on + * macOS, and an npm `bin` shim (`node_modules/.bin/distilly`) is a symlink, so a + * published CLI would silently ignore every command. + * + * Each case runs the file through a symlink and asserts it produced its real + * output, which is the difference between "did the work" and "exited cleanly". + */ + +import test from "node:test"; +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { mkdtempSync, readFileSync, rmSync, symlinkSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { fileURLToPath } from "node:url"; +import path from "node:path"; + +import { isEntryPoint } from "../src/cli/entry.mjs"; + +const root = path.dirname(path.dirname(fileURLToPath(import.meta.url))); + +/** Run `relative` through a symlink in a fresh temp directory. */ +function throughSymlink(relative, args) { + const dir = mkdtempSync(path.join(tmpdir(), "entry-")); + try { + const link = path.join(dir, path.basename(relative)); + symlinkSync(path.join(root, relative), link); + return spawnSync(process.execPath, [link, ...args], { encoding: "utf8" }); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +} + +test("the helper agrees that a symlink to this file is this file", () => { + const dir = mkdtempSync(path.join(tmpdir(), "entry-self-")); + try { + const link = path.join(dir, "entry.mjs"); + symlinkSync(fileURLToPath(import.meta.url), link); + // Simulate being started as that symlink. + const original = process.argv[1]; + process.argv[1] = link; + try { + assert.equal(isEntryPoint(import.meta.url), true); + process.argv[1] = path.join(dir, "something-else.mjs"); + assert.equal(isEntryPoint(import.meta.url), false); + process.argv[1] = path.join(dir, "does-not-exist.mjs"); + assert.equal(isEntryPoint(import.meta.url), false, "an unresolvable path is not an entry point"); + } finally { + process.argv[1] = original; + } + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); + +test("`distilly --version` works when reached through a symlink", () => { + // This is how npm invokes a package binary, so it is the shape that ships. + const result = throughSymlink(path.join("bin", "distilly.mjs"), ["--version"]); + assert.equal(result.status, 0, result.stderr); + const version = JSON.parse(readFileSync(path.join(root, "package.json"), "utf8")).version; + assert.equal(result.stdout.trim(), version, "the CLI must not exit silently under a symlink"); +}); + +test("`distilly help` works when reached through a symlink", () => { + const result = throughSymlink(path.join("bin", "distilly.mjs"), ["help"]); + assert.equal(result.status, 0, result.stderr); + assert.match(result.stdout, /用法/, "help must render, not exit 0 with nothing"); +}); + +test("`blind-test.mjs` dispatches when reached through a symlink", () => { + // Before the fix this exited 0 with no output at all: the command was dropped. + const result = throughSymlink( + path.join("scripts", "blind-test.mjs"), + ["score", "--scores", path.join(root, "no-such-scores.json")], + ); + assert.notEqual(result.status, 0, "a missing scores file is an error, not a silent success"); + assert.match( + `${result.stdout}${result.stderr}`, + /no-such-scores\.json/, + "the command must actually run and report the missing file", + ); +}); diff --git a/tests/entrypoint-docs.test.mjs b/tests/entrypoint-docs.test.mjs new file mode 100644 index 00000000..bee6d877 --- /dev/null +++ b/tests/entrypoint-docs.test.mjs @@ -0,0 +1,88 @@ +/** + * Port of the code-surface checks in `tests/test_skill_entrypoint_docs.py` + * (`test_code_uses_new_names_with_explicit_legacy_fallbacks`) plus the registry + * consistency rule from `docs/v2/ACCEPTANCE.md` phase 0. + * + * The Python original read `tools/skill_writer.py`, `tools/skill_schema.py`, + * `tools/install_codex_skill.py` and `tools/install_generated_skill_common.py`; + * those modules now live in `src/`. The documentation assertions of that suite + * (README / INSTALL / SKILL.md copy) stay in Python for ds/04-prompts and + * ds/05-agents, which own those files. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +import { PLANNED, listCommands } from "../src/commands/index.mjs"; +import { getAgent } from "../src/hosts/agents.mjs"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); +const read = (relative) => readFileSync(join(projectRoot, relative), "utf8"); + +test("the writer keeps both the new and the legacy auto-install switch", () => { + const commands = read("src/commands/skill.mjs"); + assert.match(commands, /DISTILLY_AUTO_INSTALL_CLAUDE/); + assert.match(commands, /DOT_SKILL_AUTO_INSTALL_CLAUDE/); +}); + +test("the schema still names the engine distilly in both places", () => { + const schema = read("src/skill/schema.mjs"); + assert.match(schema, /pySetDefault\(engine, "name", "distilly"\)/); + assert.match(schema, /pySetDefault\(generation, "engine", "distilly"\)/); +}); + +test("Codex keeps its documented discovery directory", () => { + assert.equal(getAgent("codex").globalPath, "~/.agents/skills/distilly"); + assert.match(read("src/install/hosts.mjs"), /\.distilly-install\.json/); +}); + +test("the collectors keep reading the private config paths, not the repo", () => { + // v2 retired `tools/*.py`, so the pre-port version of this test read files that + // no longer exist. The discipline it protected is unchanged and now lives in + // `src/collect/kit.mjs`, so it is asserted there instead of being dropped. + const kit = read(join("src", "collect", "kit.mjs")); + assert.match(kit, /join\(homedir\(\), "\.distilly"\)/, "primary credentials live under ~/.distilly"); + assert.match(kit, /"\.colleague-skill"/, "the pre-rename location stays readable"); + + // Every credentialed channel names its config file; none may read a key from + // the working tree. + for (const name of ["feishu", "dingtalk", "slack", "x", "discord", "gmail", "notion", "reddit"]) { + const source = read(join("src", "collect", `${name}.mjs`)); + assert.match(source, /export const CONFIG_FILE = "[a-z]+_config\.json";/, name); + } + + // The consent token file is a credential: it is written 0600. + assert.match(read(join("src", "consent.mjs")), /chmodSync\(staging, 0o600\)/); +}); + +test("every command the prompts may reference is registered or explicitly planned", () => { + const known = new Set([...listCommands(), ...Object.keys(PLANNED)]); + const sources = ["SKILL.md"]; + const { readdirSync, statSync } = require_fs(); + const walk = (dir) => { + for (const entry of readdirSync(join(projectRoot, dir))) { + const relative = join(dir, entry); + if (statSync(join(projectRoot, relative)).isDirectory()) walk(relative); + else if (entry.endsWith(".md")) sources.push(relative); + } + }; + walk("prompts"); + + const referenced = new Set(); + for (const source of sources) { + for (const match of read(join(source)).matchAll(/`?distilly ([a-z][a-z-]*)/g)) { + referenced.add(match[1]); + } + } + const missing = [...referenced].filter((command) => !known.has(command)).sort(); + assert.deepEqual(missing, [], `unregistered commands referenced by prompts: ${missing.join(", ")}`); +}); + +function require_fs() { + return { readdirSync: fsReaddirSync, statSync: fsStatSync }; +} + +import { readdirSync as fsReaddirSync, statSync as fsStatSync } from "node:fs"; diff --git a/tests/helpers/cli.mjs b/tests/helpers/cli.mjs new file mode 100644 index 00000000..0691a992 --- /dev/null +++ b/tests/helpers/cli.mjs @@ -0,0 +1,27 @@ +/** + * Shared CLI test helper: spawn `bin/distilly.mjs` and parse its receipt. + * Plain `.mjs` (not `*.test.mjs`) so `node --test` does not collect it twice. + */ + +import { spawnSync } from "node:child_process"; +import assert from "node:assert/strict"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +export const projectRoot = join(dirname(fileURLToPath(import.meta.url)), "..", ".."); +export const cliPath = join(projectRoot, "bin", "distilly.mjs"); + +export function runCli(args, { cwd = projectRoot, env = {} } = {}) { + return spawnSync(process.execPath, [cliPath, ...args], { + cwd, + encoding: "utf8", + env: { ...process.env, ...env }, + }); +} + +/** `--json` output is one object; everything before it is whitespace. */ +export function parseReceipt(stdout) { + const start = stdout.indexOf("{"); + assert.notEqual(start, -1, `no JSON receipt in stdout: ${stdout}`); + return JSON.parse(stdout.slice(start)); +} diff --git a/tests/install-claude-generated-skill.test.mjs b/tests/install-claude-generated-skill.test.mjs new file mode 100644 index 00000000..463e9d1c --- /dev/null +++ b/tests/install-claude-generated-skill.test.mjs @@ -0,0 +1,84 @@ +/** + * Port of `tests/test_install_claude_generated_skill.py`. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +import { installGeneratedSkillForClaude, shouldInstallCommandShim } from "../src/install/hosts.mjs"; +import { createSkill } from "../src/skill/writer.mjs"; + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-claude-")); +} + +test("install writes the Claude skill folder and its install metadata", () => { + const root = tempDir(); + try { + const generatedRoot = join(root, "skills", "celebrity"); + const claudeSkills = join(root, ".claude", "skills"); + + const skillDir = createSkill( + generatedRoot, + "zhou-qimo", + { character: "celebrity", name: "周奇墨", classification: { language: "zh-CN" } }, + "Work body", + "Persona body", + ); + + const result = installGeneratedSkillForClaude({ skillDir, skillsDir: claudeSkills, force: true }); + + const installedFile = join(claudeSkills, "celebrity-zhou-qimo", "SKILL.md"); + const metadataFile = join(claudeSkills, "celebrity-zhou-qimo", ".distilly-install.json"); + + assert.equal(result.command_name, "celebrity-zhou-qimo"); + assert.ok(existsSync(installedFile)); + assert.ok(existsSync(metadataFile)); + assert.match(readFileSync(installedFile, "utf8"), /name: celebrity-zhou-qimo/); + assert.equal(result.command_shim_installed, false); + assert.equal(result.command_path, null); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install can write the Windows command shim", () => { + const root = tempDir(); + try { + const generatedRoot = join(root, "skills", "relationship"); + const claudeSkills = join(root, ".claude", "skills"); + const claudeCommands = join(root, ".claude", "commands"); + + const skillDir = createSkill( + generatedRoot, + "mireille", + { character: "relationship", name: "Mireille" }, + "Work body", + "Persona body", + ); + + const result = installGeneratedSkillForClaude({ + skillDir, + skillsDir: claudeSkills, + commandsDir: claudeCommands, + force: true, + installCommandShim: true, + }); + + const commandFile = join(claudeCommands, "relationship-mireille.md"); + assert.equal(result.command_shim_installed, true); + assert.equal(result.command_path, commandFile); + assert.ok(existsSync(commandFile)); + assert.match(readFileSync(commandFile, "utf8"), /name: relationship-mireille/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("Windows detection only enables the command shim on Windows", () => { + assert.equal(shouldInstallCommandShim("Windows"), true); + assert.equal(shouldInstallCommandShim("Darwin"), false); +}); diff --git a/tests/install-generated-skill.test.mjs b/tests/install-generated-skill.test.mjs new file mode 100644 index 00000000..901e622a --- /dev/null +++ b/tests/install-generated-skill.test.mjs @@ -0,0 +1,135 @@ +/** + * Port of `tests/test_install_generated_skill.py` plus the drift guard that ties + * `src/install/hosts.mjs` to the shared matrix in `src/hosts/agents.mjs`. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +import { + defaultSkillsDir, + generatedSkillsRoot, + installGeneratedSkill, + repoInstallDir, + supportedHosts, +} from "../src/install/hosts.mjs"; +import { listAgents } from "../src/hosts/agents.mjs"; +import { createSkill } from "../src/skill/writer.mjs"; + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-install-")); +} + +function createLegacySkill(root) { + const skillDir = createSkill( + root, + "mireille", + { character: "relationship", name: "Mireille" }, + "Work body", + "Persona body", + ); + const skillPath = join(skillDir, "SKILL.md"); + writeFileSync( + skillPath, + readFileSync(skillPath, "utf8").replace("name: relationship-mireille", "name: relationship_mireille"), + "utf8", + ); + const metaPath = join(skillDir, "meta.json"); + const meta = JSON.parse(readFileSync(metaPath, "utf8")); + delete meta.artifacts; + writeFileSync(metaPath, JSON.stringify(meta), "utf8"); + return skillDir; +} + +test("the installer host list is exactly the shared coding-agent matrix", () => { + assert.deepEqual([...supportedHosts()].sort(), [...listAgents()].sort()); + for (const host of listAgents()) { + assert.equal(typeof repoInstallDir(host, { home: "/h", env: {} }), "string"); + assert.equal(typeof generatedSkillsRoot(host, { home: "/h", env: {} }), "string"); + } +}); + +test("default skills dirs cover all documented hosts", () => { + const home = "/example/home"; + const expected = { + "claude-code": join(home, ".claude", "skills"), + openclaw: join(home, ".openclaw", "workspace", "skills"), + hermes: join(home, ".hermes", "skills", "distilly-generated"), + codex: join(home, ".agents", "skills"), + "deepseek-harness": join(home, ".dsh", "skills"), + pi: join(home, ".pi", "agent", "skills"), + "grok-build": join(home, ".grok", "skills"), + opencode: join(home, ".config", "opencode", "skills"), + }; + for (const [host, dir] of Object.entries(expected)) { + assert.equal(defaultSkillsDir(host, { home, env: {} }), dir, host); + } + assert.equal( + defaultSkillsDir("deepseek-harness", { home, env: { DSH_HOME: "/custom/dsh" } }), + join("/custom/dsh", "skills"), + ); +}); + +test("install rewrites only the legacy copy to the canonical name", () => { + const root = tempDir(); + try { + const source = createLegacySkill(join(root, "generated")); + const skillsDir = join(root, ".hermes", "skills", "distilly-generated"); + + const result = installGeneratedSkill({ skillDir: source, skillsDir, force: true, host: "hermes" }); + + const installed = result.skill_dir; + assert.equal(installed.split("/").pop(), "relationship-mireille"); + assert.deepEqual(readdirSync(installed).sort(), [".distilly-install.json", "SKILL.md"]); + assert.match(readFileSync(join(installed, "SKILL.md"), "utf8"), /name: relationship-mireille/); + assert.match(readFileSync(join(source, "SKILL.md"), "utf8"), /name: relationship_mireille/); + + const record = JSON.parse(readFileSync(join(installed, ".distilly-install.json"), "utf8")); + assert.equal(record.host, "hermes"); + assert.equal(record.command_name, "relationship-mireille"); + assert.equal(record.slug, "mireille"); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install rejects ancestor or descendant destinations", () => { + const root = tempDir(); + try { + const source = createLegacySkill(join(root, "generated")); + + assert.throws( + () => installGeneratedSkill({ skillDir: source, skillsDir: source, force: true, host: "test" }), + /must not overlap/, + ); + + const ancestorInstall = join(root, "host", "relationship-bundle"); + mkdirSync(ancestorInstall, { recursive: true }); + const nestedSource = join(ancestorInstall, "bundle"); + const { renameSync } = await_rename(); + renameSync(source, nestedSource); + assert.throws( + () => + installGeneratedSkill({ + skillDir: nestedSource, + skillsDir: join(root, "host"), + force: true, + host: "test", + }), + /must not overlap/, + ); + + assert.ok(readFileSync(join(nestedSource, "meta.json"), "utf8")); + assert.ok(readFileSync(join(nestedSource, "work.md"), "utf8")); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +// `renameSync` is imported lazily so the test file keeps a single import style. +function await_rename() { + return { renameSync: (from, to) => import("node:fs").then(() => from && to) && renameSyncImpl(from, to) }; +} diff --git a/tests/install-hermes-skill.test.mjs b/tests/install-hermes-skill.test.mjs new file mode 100644 index 00000000..d81592e6 --- /dev/null +++ b/tests/install-hermes-skill.test.mjs @@ -0,0 +1,116 @@ +/** + * Port of `tests/test_install_hermes_skill.py` (the repo-level installer that + * `install_codex_skill.py` and `install_openclaw_skill.py` also implement). + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +import { installRepoSkill } from "../src/install/hosts.mjs"; + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-hermes-")); +} + +test("install copies the repo layout", () => { + const root = tempDir(); + try { + const source = join(root, "source"); + const destination = join(root, "dest", "distilly"); + mkdirSync(source); + writeFileSync(join(source, "SKILL.md"), "name: distilly\n", "utf8"); + writeFileSync(join(source, "README.md"), "# Distilly\n", "utf8"); + + const installed = installRepoSkill({ source, destination }); + assert.equal(installed, destination); + assert.ok(existsSync(join(destination, "SKILL.md"))); + assert.ok(existsSync(join(destination, "README.md"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install skips metadata a host should not receive", () => { + const root = tempDir(); + try { + const source = join(root, "source"); + mkdirSync(join(source, "__pycache__"), { recursive: true }); + mkdirSync(join(source, ".git"), { recursive: true }); + writeFileSync(join(source, "SKILL.md"), "name: distilly\n", "utf8"); + writeFileSync(join(source, "__pycache__", "x.pyc"), "", "utf8"); + writeFileSync(join(source, ".git", "config"), "", "utf8"); + writeFileSync(join(source, "stale.pyc"), "", "utf8"); + + const destination = join(root, "dest", "distilly"); + installRepoSkill({ source, destination }); + + assert.ok(existsSync(join(destination, "SKILL.md"))); + assert.equal(existsSync(join(destination, "__pycache__")), false); + assert.equal(existsSync(join(destination, ".git")), false); + assert.equal(existsSync(join(destination, "stale.pyc")), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("dry run does not write and tolerates an existing destination", () => { + const root = tempDir(); + try { + const source = join(root, "source"); + const destination = join(root, "dest", "distilly"); + mkdirSync(source); + writeFileSync(join(source, "SKILL.md"), "name: distilly\n", "utf8"); + + installRepoSkill({ source, destination, dryRun: true }); + assert.equal(existsSync(destination), false); + + mkdirSync(destination, { recursive: true }); + const result = installRepoSkill({ source, destination, dryRun: true }); + assert.equal(result, destination); + assert.ok(existsSync(destination)); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("dry run does not delete the source when it is already the destination", () => { + const root = tempDir(); + try { + const source = join(root, "distilly"); + mkdirSync(source); + const skillFile = join(source, "SKILL.md"); + writeFileSync(skillFile, "name: distilly\n", "utf8"); + + const result = installRepoSkill({ source, destination: source, force: true }); + + assert.equal(result, source); + assert.ok(existsSync(skillFile)); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("install rejects nested or ancestor destinations", () => { + const root = tempDir(); + try { + const source = join(root, "parent", "source"); + mkdirSync(source, { recursive: true }); + const skillFile = join(source, "SKILL.md"); + writeFileSync(skillFile, "name: distilly\n", "utf8"); + const nested = join(source, ".hermes", "skills", "distilly"); + + assert.throws(() => installRepoSkill({ source, destination: nested, force: true }), /must not overlap/); + assert.throws( + () => installRepoSkill({ source, destination: join(root, "parent"), force: true }), + /must not overlap/, + ); + + assert.ok(existsSync(skillFile)); + assert.equal(existsSync(nested), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); diff --git a/tests/install-openclaw-and-codex.test.mjs b/tests/install-openclaw-and-codex.test.mjs new file mode 100644 index 00000000..0aa89332 --- /dev/null +++ b/tests/install-openclaw-and-codex.test.mjs @@ -0,0 +1,157 @@ +/** + * Port of `tests/test_install_openclaw_and_codex.py`: the repo-level installers + * share one implementation, the generated-skill installers differ only in the + * host id they record. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; + +import { installGeneratedSkill, installRepoSkill } from "../src/install/hosts.mjs"; +import { createSkill } from "../src/skill/writer.mjs"; + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-openclaw-")); +} + +test("openclaw and codex repo installers copy the repo layout", () => { + const root = tempDir(); + try { + const source = join(root, "source"); + mkdirSync(source); + writeFileSync(join(source, "SKILL.md"), "name: distilly\n", "utf8"); + writeFileSync(join(source, "README.md"), "# Distilly\n", "utf8"); + + const openclawDest = join(root, "openclaw", "distilly"); + const codexDest = join(root, "agents", "distilly"); + + assert.equal(installRepoSkill({ source, destination: openclawDest }), openclawDest); + assert.equal(installRepoSkill({ source, destination: codexDest }), codexDest); + assert.ok(existsSync(join(openclawDest, "SKILL.md"))); + assert.ok(existsSync(join(codexDest, "SKILL.md"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("repo installers do not delete the source when it is already the destination", () => { + const root = tempDir(); + try { + const source = join(root, "distilly"); + mkdirSync(source); + const skillFile = join(source, "SKILL.md"); + writeFileSync(skillFile, "name: distilly\n", "utf8"); + + assert.equal(installRepoSkill({ source, destination: source, force: true }), source); + assert.ok(existsSync(skillFile)); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("repo installers reject nested or ancestor destinations", () => { + const root = tempDir(); + try { + const source = join(root, "parent", "source"); + mkdirSync(source, { recursive: true }); + const skillFile = join(source, "SKILL.md"); + writeFileSync(skillFile, "name: distilly\n", "utf8"); + const nested = join(source, ".agents", "skills", "distilly"); + + assert.throws(() => installRepoSkill({ source, destination: nested, force: true }), /must not overlap/); + assert.throws( + () => installRepoSkill({ source, destination: join(root, "parent"), force: true }), + /must not overlap/, + ); + + assert.ok(existsSync(skillFile)); + assert.equal(existsSync(nested), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("openclaw generated-skill installer writes the host skill folder", () => { + const root = tempDir(); + try { + const generatedRoot = join(root, "skills", "relationship"); + const openclawSkills = join(root, ".openclaw", "workspace", "skills"); + + const skillDir = createSkill( + generatedRoot, + "mireille", + { character: "relationship", name: "Mireille" }, + "Work body", + "Persona body", + ); + + const result = installGeneratedSkill({ + skillDir, + skillsDir: openclawSkills, + force: true, + host: "openclaw", + }); + + const installedFile = join(openclawSkills, "relationship-mireille", "SKILL.md"); + const metadataFile = join(openclawSkills, "relationship-mireille", ".distilly-install.json"); + + assert.equal(result.command_name, "relationship-mireille"); + assert.ok(existsSync(installedFile)); + assert.ok(existsSync(metadataFile)); + assert.match(readFileSync(installedFile, "utf8"), /name: relationship-mireille/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("codex generated-skill installer writes the host skill folder", () => { + const root = tempDir(); + try { + const generatedRoot = join(root, "skills", "celebrity"); + const codexSkills = join(root, ".agents", "skills"); + + const skillDir = createSkill( + generatedRoot, + "zhou-qimo", + { character: "celebrity", name: "周奇墨", classification: { language: "zh-CN" } }, + "Work body", + "Persona body", + ); + + const result = installGeneratedSkill({ skillDir, skillsDir: codexSkills, force: true, host: "codex" }); + + const installedFile = join(codexSkills, "celebrity-zhou-qimo", "SKILL.md"); + const metadataFile = join(codexSkills, "celebrity-zhou-qimo", ".distilly-install.json"); + + assert.equal(result.command_name, "celebrity-zhou-qimo"); + assert.ok(existsSync(installedFile)); + assert.ok(existsSync(metadataFile)); + assert.match(readFileSync(installedFile, "utf8"), /name: celebrity-zhou-qimo/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("an existing install is only replaced with force", () => { + const root = tempDir(); + try { + const generatedRoot = join(root, "skills", "relationship"); + const skillsDir = join(root, ".agents", "skills"); + const skillDir = createSkill( + generatedRoot, + "mireille", + { character: "relationship", name: "Mireille" }, + "Work body", + "Persona body", + ); + + installGeneratedSkill({ skillDir, skillsDir, host: "codex" }); + assert.throws(() => installGeneratedSkill({ skillDir, skillsDir, host: "codex" }), /already exists/); + installGeneratedSkill({ skillDir, skillsDir, host: "codex", force: true }); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); diff --git a/tests/package-payload.test.mjs b/tests/package-payload.test.mjs new file mode 100644 index 00000000..e42a7e31 --- /dev/null +++ b/tests/package-payload.test.mjs @@ -0,0 +1,147 @@ +/** + * The published package has to be runnable, and the manifest is what decides that. + * + * `npm publish` ships exactly `package.json`'s `files` (plus a few npm always + * includes), so a manifest that drifts from the tree produces a tarball that + * installs and then dies on `import "../src/cli/args.mjs"`. That is not + * hypothetical: the v2 port removed `tools/` and `requirements.txt` and added + * `src/` and `assets/`, but `files` kept listing the first two and never gained + * the second two. Every gate stayed green — `--check-package` validated the + * *working tree*, which had everything — while the extracted tarball could not + * even print `--version`. + * + * These tests drive `validatePayload` against throwaway roots, so each failure + * mode is named precisely and costs no `npm pack`. `scripts/check_release.mjs` + * runs the real pack-and-run. + */ + +import test from "node:test"; +import assert from "node:assert/strict"; +import { mkdtempSync, readFileSync, rmSync, symlinkSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { fileURLToPath } from "node:url"; +import path from "node:path"; + +import { payloadEntries, validatePayload } from "../bin/distilly.mjs"; + +const root = path.dirname(path.dirname(fileURLToPath(import.meta.url))); +const manifest = JSON.parse(readFileSync(path.join(root, "package.json"), "utf8")); + +/** + * A root that really has every `payloadEntries` path (symlinked, so the check + * sees them) but the candidate `files` list under test. + * + * `package.json` itself is written as a real file and never symlinked: writing + * through a symlink to the repository's manifest would edit the repository — the + * first version of this helper did exactly that and blanked the real `files`. + */ +function rootWith(files) { + const dir = mkdtempSync(path.join(tmpdir(), "payload-")); + for (const entry of payloadEntries) { + if (entry === "package.json") continue; + symlinkSync(path.join(root, entry), path.join(dir, entry)); + } + writeFileSync(path.join(dir, "package.json"), JSON.stringify({ ...manifest, files }, null, 2)); + return dir; +} + +const withManifest = (files, body) => { + const dir = rootWith(files); + try { + return body(dir); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}; + +test("the manifest ships every path the runtime imports", () => { + for (const required of ["bin/", "src/", "assets/", "scripts/", "SKILL.md", "prompts/"]) { + assert.ok(manifest.files.includes(required), `package.json \`files\` must list ${required}`); + } +}); + +test("the shipped manifest passes its own gate", () => { + withManifest(manifest.files, (dir) => assert.doesNotThrow(() => validatePayload(dir))); +}); + +test("a manifest that drops src/ is rejected, not packed", () => { + const files = manifest.files.filter((entry) => entry !== "src/"); + withManifest(files, (dir) => { + assert.throws( + () => validatePayload(dir), + (error) => /would omit required paths: src/.test(error.message) && /add "src\/"/.test(error.remedy), + "omitting src/ must fail the prepack gate with the fix in the remedy", + ); + }); +}); + +test("a manifest that drops assets/ is rejected, not packed", () => { + const files = manifest.files.filter((entry) => entry !== "assets/"); + withManifest(files, (dir) => { + assert.throws(() => validatePayload(dir), /would omit required paths: assets/); + }); +}); + +test("a directory pattern covers everything below it", () => { + // `src/` must be understood as "all of src", not as the literal name `src`. + assert.ok(manifest.files.includes("src/")); + withManifest(["src/"], (dir) => { + assert.throws(() => validatePayload(dir), /would omit required paths: (?!src)/, "src/ alone must satisfy the src entry"); + }); +}); + +test("stale Python-era manifest entries are rejected", () => { + // `tools/` no longer exists, so listing it makes the manifest a lie about what + // the package contains — this is the other half of the shipped bug. + withManifest([...manifest.files, "tools/"], (dir) => { + assert.throws(() => validatePayload(dir), /names paths that do not exist: tools/); + }); +}); + +test("a manifest with no `files` at all is rejected", () => { + withManifest([], (dir) => { + assert.throws(() => validatePayload(dir), /declares no `files`/); + }); +}); + +test("these fixtures never write through to the repository manifest", () => { + // Regression guard for the helper itself: symlinking package.json and then + // writing the candidate manifest overwrote the real one (files became []). + const manifestPath = path.join(root, "package.json"); + const before = readFileSync(manifestPath, "utf8"); + withManifest([], (dir) => assert.throws(() => validatePayload(dir))); + withManifest(manifest.files.filter((entry) => entry !== "src/"), (dir) => assert.throws(() => validatePayload(dir))); + assert.equal(readFileSync(manifestPath, "utf8"), before, "the repository manifest must be untouched"); +}); + +test("`npm test` and CI run the same command, and it names the tests directory", () => { + // A bare `node --test` also collects `scripts/blind-test.mjs` (`**/*-test.mjs`) + // and records its usage error as a failing test — which is how CI went red + // while every local command looked green. + const ci = readFileSync(path.join(root, ".github", "workflows", "ci.yml"), "utf8"); + assert.match(ci, /^\s*run: npm test\s*$/m, "CI must invoke `npm test`"); + assert.equal(manifest.scripts.test, 'node --test "tests/*.test.mjs"'); + assert.equal( + /\bnode --test\s*$/.test(manifest.scripts.test), + false, + "`npm test` must not be a bare `node --test`: it would collect scripts/ as tests", + ); +}); + +test("`engines` matches what CI actually tests", () => { + const ci = readFileSync(path.join(root, ".github", "workflows", "ci.yml"), "utf8"); + const matrix = /node-version:\s*\[([^\]]+)\]/.exec(ci)?.[1] ?? ""; + const versions = matrix.split(",").map((entry) => entry.trim().replace(/"/g, "")); + const minimum = Number(manifest.engines.node.replace(/[^\d]/g, "")); + assert.ok(versions.length > 0, "CI must declare a node-version matrix"); + for (const version of versions) { + assert.ok( + Number(version) >= minimum, + `CI tests Node ${version} but the package claims to need >= ${minimum}`, + ); + } + assert.ok( + versions.includes(String(minimum)), + `the oldest supported Node (${minimum}) is not in the CI matrix, so the floor is untested`, + ); +}); diff --git a/tests/release-check.test.mjs b/tests/release-check.test.mjs new file mode 100644 index 00000000..feb23e1e --- /dev/null +++ b/tests/release-check.test.mjs @@ -0,0 +1,37 @@ +/** + * The release check must pass, and must be able to fail. + * + * `scripts/check_release.mjs` covers what the other gates do not: the version every + * surface prints, the schema marker, the installer's carried directories, generated + * artefacts, and the gates/CI/docs a release claims. + */ + +import test from "node:test"; +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { join } from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = fileURLToPath(new URL("..", import.meta.url)); +const SCRIPT = join(HERE, "scripts", "check_release.mjs"); + +function run(...args) { + const result = spawnSync(process.execPath, [SCRIPT, "--json", ...args], { cwd: HERE, encoding: "utf8" }); + return { status: result.status, report: JSON.parse(result.stdout.slice(result.stdout.indexOf("{"))) }; +} + +test("every release-time check passes", () => { + const { status, report } = run(); + assert.equal(status, 0, JSON.stringify(report.rows.filter((row) => !row.ok), null, 2)); + assert.equal(report.ok, true); + assert.ok(report.rows.length >= 7); + for (const row of report.rows) assert.ok(row.evidence.length > 0, `${row.name} must say what it read`); +}); + +test("a mismatched release tag fails the run", () => { + const { status, report } = run("--tag", "v9.9.9"); + assert.equal(status, 1, "a tag that disagrees with package.json must fail"); + const tagRow = report.rows.find((row) => row.name.includes("tag")); + assert.equal(tagRow.ok, false); + assert.match(tagRow.evidence, /v9\.9\.9/); +}); diff --git a/tests/schema-migration.test.mjs b/tests/schema-migration.test.mjs new file mode 100644 index 00000000..f8fe0163 --- /dev/null +++ b/tests/schema-migration.test.mjs @@ -0,0 +1,165 @@ +/** + * Schema v4 migration: additive, idempotent, never rewrites a body. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { createHash } from "node:crypto"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, readdirSync, rmSync, statSync, writeFileSync } from "node:fs"; +import { spawnSync } from "node:child_process"; +import { tmpdir } from "node:os"; +import { join, relative } from "node:path"; + +import { SCHEMA_VERSION } from "../src/skill/schema.mjs"; +import { V4_DIRECTORIES, findSkillDirs, migrateSkillDir, readSchemaVersion } from "../src/skill/migrate.mjs"; + +const BIN = join(import.meta.dirname, "..", "bin", "distilly.mjs"); +const PERSONA = "# 张三\n\n先说结论,再讲为什么。\n"; +const SKILL = "---\nname: 张三\n---\n\n# 张三\n"; + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-migrate-")); +} + +/** A minimal v3 directory: six artifacts minus the v4 layout. */ +function makeV3(baseDir, slug = "old-person") { + const dir = join(baseDir, "skills", "colleague", slug); + mkdirSync(dir, { recursive: true }); + writeFileSync(join(dir, "SKILL.md"), SKILL, "utf8"); + writeFileSync(join(dir, "persona.md"), PERSONA, "utf8"); + writeFileSync(join(dir, "work.md"), "# 工作\n", "utf8"); + writeFileSync( + join(dir, "meta.json"), + JSON.stringify({ name: "张三", slug, character: "colleague", schema_version: "3", lifecycle: { version: "v1" } }, null, 2), + "utf8", + ); + writeFileSync(join(dir, "manifest.json"), JSON.stringify({ slug, schema_version: "3", artifacts: {} }, null, 2), "utf8"); + return dir; +} + +/** Every file under `dir` with its sha256, for byte-equality assertions. */ +function fingerprint(dir) { + const out = {}; + const walk = (current) => { + for (const entry of readdirSync(current, { withFileTypes: true }).sort((a, b) => a.name.localeCompare(b.name))) { + const full = join(current, entry.name); + if (entry.isDirectory()) walk(full); + else out[relative(dir, full)] = createHash("sha256").update(readFileSync(full)).digest("hex"); + } + }; + walk(dir); + return out; +} + +test("the engine reports schema v4", () => { + assert.equal(SCHEMA_VERSION, "4"); +}); + +test("migrating a v3 directory adds the v4 layout and touches no body", () => { + const base = tempDir(); + try { + const dir = makeV3(base); + const before = { persona: readFileSync(join(dir, "persona.md")), skill: readFileSync(join(dir, "SKILL.md")) }; + + assert.equal(readSchemaVersion(dir), "3"); + const result = migrateSkillDir(dir); + + assert.equal(result.from, "3"); + assert.equal(result.to, "4"); + assert.equal(result.changed, true); + for (const relativePath of V4_DIRECTORIES) { + assert.ok(existsSync(join(dir, relativePath)), `${relativePath} must exist after migration`); + } + assert.deepEqual(JSON.parse(readFileSync(join(dir, "knowledge", "index.json"), "utf8")), []); + + assert.equal(JSON.parse(readFileSync(join(dir, "meta.json"), "utf8")).schema_version, "4"); + assert.equal(JSON.parse(readFileSync(join(dir, "manifest.json"), "utf8")).schema_version, "4"); + assert.ok(readFileSync(join(dir, "persona.md")).equals(before.persona), "persona.md must keep its bytes"); + assert.ok(readFileSync(join(dir, "SKILL.md")).equals(before.skill), "SKILL.md must keep its bytes"); + assert.equal(readSchemaVersion(dir), "4"); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); + +test("running the migration twice writes nothing the second time", () => { + const base = tempDir(); + try { + const dir = makeV3(base); + migrateSkillDir(dir); + const after = fingerprint(dir); + const stamps = Object.fromEntries(Object.keys(after).map((key) => [key, statSync(join(dir, key)).mtimeMs])); + + const second = migrateSkillDir(dir); + + assert.equal(second.changed, false); + assert.deepEqual(second.actions, []); + assert.deepEqual(fingerprint(dir), after, "no file may change on a second run"); + for (const [key, mtime] of Object.entries(stamps)) { + assert.equal(statSync(join(dir, key)).mtimeMs, mtime, `${key} was rewritten`); + } + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); + +test("--dry-run reports the plan and writes nothing", () => { + const base = tempDir(); + try { + const dir = makeV3(base); + const before = fingerprint(dir); + + const result = migrateSkillDir(dir, { dryRun: true }); + + assert.equal(result.changed, true); + assert.ok(result.actions.some((action) => action.includes("knowledge/raw"))); + assert.deepEqual(fingerprint(dir), before, "a dry run must not touch the directory"); + assert.equal(readSchemaVersion(dir), "3"); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); + +test("findSkillDirs walks families and skips directories without artifacts", () => { + const base = tempDir(); + try { + makeV3(base, "one"); + makeV3(base, "two"); + mkdirSync(join(base, "skills", "colleague", "not-a-skill"), { recursive: true }); + mkdirSync(join(base, "skills", "relationship", "three"), { recursive: true }); + writeFileSync(join(base, "skills", "relationship", "three", "meta.json"), JSON.stringify({ slug: "three" }), "utf8"); + + const found = findSkillDirs(base).map((dir) => relative(base, dir).replaceAll("\\", "/")); + + assert.deepEqual(found, ["skills/colleague/one", "skills/colleague/two", "skills/relationship/three"]); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); + +test("the CLI reports a contract-shaped receipt and is idempotent", () => { + const base = tempDir(); + try { + makeV3(base); + const run = () => { + const result = spawnSync(process.execPath, [BIN, "skill", "migrate", "--base-dir", base, "--json"], { + encoding: "utf8", + }); + assert.equal(result.status, 0, result.stderr); + return JSON.parse(result.stdout.slice(result.stdout.indexOf("{"))); + }; + + const first = run(); + assert.equal(first.command, "skill migrate"); + assert.equal(first.ok, true); + assert.equal(first.target_schema_version, "4"); + assert.equal(first.skills.length, 1); + assert.equal(first.skills[0].from, "3"); + assert.equal(first.skills[0].changed, true); + + const second = run(); + assert.equal(second.skills[0].changed, false, "a migrated directory must report no change"); + } finally { + rmSync(base, { recursive: true, force: true }); + } +}); diff --git a/tests/skill-writer.test.mjs b/tests/skill-writer.test.mjs new file mode 100644 index 00000000..a02ecd3d --- /dev/null +++ b/tests/skill-writer.test.mjs @@ -0,0 +1,574 @@ +/** + * Port of `tests/test_skill_writer.py` (SkillWriterTest, VersionManagerTest and + * PromptPresetTest), asserting the same behaviour against `src/skill/*`. + */ + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, renameSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import { fileURLToPath } from "node:url"; + +import { + WORK_ONLY_FALLBACK_EN, + WORK_ONLY_FALLBACK_ZH, + createSkill, + listSkills, + slugify, + updateSkill, +} from "../src/skill/writer.mjs"; +import { SCHEMA_VERSION } from "../src/skill/schema.mjs"; +import { PLANNED, listCommands } from "../src/commands/index.mjs"; +import { backupCurrentVersion, rollback } from "../src/skill/versions.mjs"; +import { getCharacterPreset, getResearchProfilePreset, resolveExistingStorageRoot } from "../src/skill/presets.mjs"; +import { validatePathSegment } from "../src/skill/schema.mjs"; + +const projectRoot = join(dirname(fileURLToPath(import.meta.url)), ".."); + +function tempDir() { + return mkdtempSync(join(tmpdir(), "dst-writer-")); +} + +function readJson(path) { + return JSON.parse(readFileSync(path, "utf8")); +} + +test("slugify produces portable kebab-case", () => { + assert.equal(slugify("Zadie Smith"), "zadie-smith"); + assert.equal(slugify("Élodie"), "elodie"); + assert.equal(slugify("A/B"), "a-b"); +}); + +test("createSkill rejects an unsafe slug before writing", () => { + const root = tempDir(); + try { + assert.throws( + () => createSkill(join(root, "skills", "colleague"), "../escape", { name: "Unsafe" }, "Work body", "Persona body"), + /kebab-case/, + ); + assert.equal(existsSync(join(root, "skills", "escape")), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("legacy path segments are Windows safe", () => { + assert.equal(validatePathSegment("Zadie Smith"), "Zadie Smith"); + assert.equal(validatePathSegment("Élodie"), "Élodie"); + for (const value of ["C:", "foo:bar", "CON", "nul.txt", "trailing.", "trailing "]) { + assert.throws(() => validatePathSegment(value), /safe path segment/, value); + } +}); + +test("create colleague uses portable names and adds the engine schema", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const meta = { + name: "Eulalie", + profile: { company: "ByteDance", level: "L2-1", role: "Backend Engineer", mbti: "INTJ" }, + tags: { personality: ["direct", "data-driven"], culture: ["byte-dance-style"] }, + }; + + const skillDir = createSkill(baseDir, "zhangsan", meta, "Work body", "Persona body"); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const manifest = readJson(join(skillDir, "manifest.json")); + const combinedSkill = readFileSync(join(skillDir, "SKILL.md"), "utf8"); + const workSkill = readFileSync(join(skillDir, "work_skill.md"), "utf8"); + const personaSkill = readFileSync(join(skillDir, "persona_skill.md"), "utf8"); + + assert.equal(savedMeta.schema_version, SCHEMA_VERSION, "the engine stamps its current schema version"); + assert.equal(savedMeta.kind, "meta-skill"); + assert.equal(savedMeta.character, "colleague"); + assert.equal(savedMeta.preset, "distilly.colleague.v1"); + assert.equal(savedMeta.engine.name, "distilly"); + assert.equal(savedMeta.generation.engine, "distilly"); + assert.equal(savedMeta.type, "colleague"); + assert.equal(savedMeta.id, "meta-skill.colleague.zhangsan"); + assert.equal(savedMeta.artifacts.combined_name, "colleague-zhangsan"); + assert.equal(savedMeta.artifacts.combined_command, "colleague-zhangsan"); + assert.equal(savedMeta.compat.legacy_command, "/create-colleague"); + assert.equal(manifest.kind, "meta-skill"); + assert.equal(manifest.character, "colleague"); + assert.equal(manifest.preset, "distilly.colleague.v1"); + assert.equal(manifest.install.slash_commands.default, "colleague-zhangsan"); + assert.deepEqual(manifest.install.compatible_runtimes, [ + "claude-code", + "openclaw", + "hermes", + "codex", + "deepseek-harness", + "grok-build", + "pi", + "opencode", + ]); + assert.equal(manifest.install.installers.openclaw, "tools/install_openclaw_generated_skill.py"); + assert.equal(manifest.install.installers.codex, "tools/install_codex_generated_skill.py"); + assert.match(combinedSkill, /name: colleague-zhangsan/); + assert.match(combinedSkill, /## PART A: Work/); + assert.match(workSkill, /name: colleague-zhangsan-work/); + assert.match(workSkill, /work capability only/); + assert.match(personaSkill, /name: colleague-zhangsan-persona/); + assert.match(personaSkill, /persona only/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create relationship uses the character preset metadata", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "relationship"); + const skillDir = createSkill(baseDir, "mireille", { character: "relationship", name: "Mireille", profile: { role: "Designer" } }, "Work body", "Persona body"); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const manifest = readJson(join(skillDir, "manifest.json")); + const combinedSkill = readFileSync(join(skillDir, "SKILL.md"), "utf8"); + + assert.equal(savedMeta.kind, "meta-skill"); + assert.equal(savedMeta.character, "relationship"); + assert.equal(savedMeta.preset, "distilly.relationship.v1"); + assert.equal(savedMeta.type, "relationship"); + assert.equal(savedMeta.classification.gallery_category, "Relationship"); + assert.equal(savedMeta.compat.legacy_storage_root, "skills/relationship"); + assert.equal(manifest.id, "meta-skill.relationship.mireille"); + assert.equal(manifest.character, "relationship"); + assert.equal(savedMeta.artifacts.combined_command, "relationship-mireille"); + assert.match(combinedSkill, /name: relationship-mireille/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create renders Chinese chrome when the language is zh-CN", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "relationship"); + const skillDir = createSkill( + baseDir, + "mireille", + { character: "relationship", name: "Mireille", classification: { language: "zh-CN" } }, + "Work body", + "Persona body", + ); + + assert.match(readFileSync(join(skillDir, "SKILL.md"), "utf8"), /## PART A:工作能力/); + assert.match(readFileSync(join(skillDir, "SKILL.md"), "utf8"), /运行规则/); + assert.match(readFileSync(join(skillDir, "work_skill.md"), "utf8"), /仅 Work,无 Persona/); + assert.match(readFileSync(join(skillDir, "persona_skill.md"), "utf8"), /仅 Persona,无工作能力/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("the work-only skill replaces the persona handoff", () => { + const zhHandoff = "如果被问到职责范围外的问题,以该同事的方式回应(参见 Persona 部分)。"; + const enHandoff = + "If you are asked a question outside your recorded responsibilities, respond in this colleague's style (see the Persona section)."; + const zhWorkContent = `## 工作能力使用说明\n\n当用户要求你完成以下任务时,严格按照上述规范执行。\n\n${zhHandoff}\n`; + const enWorkContent = `## Scope rule\n\nIf asked outside your recorded responsibilities:\n- State the evidence gap\n\n## Persona naming note\n\nKeep this documentation sentence.\n\n${enHandoff}\n`; + + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const shared = { company: "ByteDance", level: "L2-1", role: "Backend Engineer" }; + + const zhDir = createSkill(join(baseDir, "zh"), "zhangsan", { name: "Eulalie", language: "zh-CN", profile: shared }, zhWorkContent, "Persona body"); + const enDir = createSkill(join(baseDir, "en"), "zhangsan", { name: "Eulalie", language: "en", profile: shared }, enWorkContent, "Persona body"); + + const zhStoredWork = readFileSync(join(zhDir, "work.md"), "utf8"); + const zhCombined = readFileSync(join(zhDir, "SKILL.md"), "utf8"); + const enStoredWork = readFileSync(join(enDir, "work.md"), "utf8"); + const enCombined = readFileSync(join(enDir, "SKILL.md"), "utf8"); + const zhWorkSkill = readFileSync(join(zhDir, "work_skill.md"), "utf8"); + const enWorkSkill = readFileSync(join(enDir, "work_skill.md"), "utf8"); + + assert.match(zhStoredWork, new RegExp(zhHandoff)); + assert.match(zhCombined, new RegExp(zhHandoff)); + assert.match(enStoredWork, new RegExp(enHandoff)); + assert.match(enCombined, new RegExp(enHandoff)); + assert.equal(zhWorkSkill.includes(zhHandoff), false); + assert.equal(enWorkSkill.includes(enHandoff), false); + assert.match(enWorkSkill, /If asked outside your recorded responsibilities:/); + assert.match(enWorkSkill, /## Persona naming note/); + assert.match(enWorkSkill, /Keep this documentation sentence\./); + assert.ok(zhWorkSkill.includes(WORK_ONLY_FALLBACK_ZH)); + assert.ok(enWorkSkill.includes(WORK_ONLY_FALLBACK_EN)); + assert.match(zhWorkSkill, /不要臆造缺失信息/); + assert.equal(zhWorkSkill.includes("不要推断"), false); + assert.match(enWorkSkill, /Do not fabricate missing information/); + assert.equal(enWorkSkill.includes("Do not infer"), false); + assert.equal(zhCombined.includes(WORK_ONLY_FALLBACK_ZH), false); + assert.equal(enCombined.includes(WORK_ONLY_FALLBACK_EN), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create celebrity adds research dirs and the toolchain", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + const skillDir = createSkill( + baseDir, + "zadie-smith", + { + character: "celebrity", + name: "Zadie Smith", + profile: { identity: "Novelist", known_for: "Essay and criticism" }, + tags: ["literature", "essay", "public-intellectual"], + knowledge_sources: ["interview", "essay"], + }, + "Work body", + "Persona body", + ); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const manifest = readJson(join(skillDir, "manifest.json")); + + assert.equal(savedMeta.character, "celebrity"); + assert.equal(savedMeta.preset, "distilly.celebrity.v1"); + assert.equal(savedMeta.research_profile, "budget-friendly"); + assert.ok("research_tools" in savedMeta.engine); + assert.equal(savedMeta.engine.research_profile, "budget-friendly"); + assert.ok("research_tools" in manifest.toolchain); + assert.equal(manifest.research_profile, "budget-friendly"); + assert.deepEqual(savedMeta.classification.tags, ["literature", "essay", "public-intellectual"]); + assert.match(savedMeta.summary, /Novelist/); + assert.match(savedMeta.summary, /Essay and criticism/); + assert.ok(existsSync(join(skillDir, "knowledge", "research", "raw"))); + assert.ok(existsSync(join(skillDir, "knowledge", "research", "merged"))); + assert.ok(existsSync(join(skillDir, "knowledge", "transcripts"))); + assert.ok(existsSync(join(skillDir, "knowledge", "subtitles"))); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create celebrity budget-unfriendly embeds the profile config", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + const skillDir = createSkill( + baseDir, + "xu-zhisheng", + { character: "celebrity", research_profile: "budget-unfriendly", name: "Xu Zhisheng", classification: { language: "zh-CN" } }, + "Work body", + "Persona body", + ); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const manifest = readJson(join(skillDir, "manifest.json")); + + assert.equal(savedMeta.research_profile, "budget-unfriendly"); + assert.equal(savedMeta.engine.quality_profile, "budget-unfriendly"); + assert.ok(Object.values(savedMeta.engine.research_profile_bundle).includes("prompts/celebrity/budget_unfriendly/research.md")); + assert.ok(Object.values(savedMeta.engine.research_profile_bundle).includes("prompts/celebrity/budget_unfriendly/audit.md")); + assert.ok(savedMeta.engine.research_profile_references.includes("references/celebrity_budget_unfriendly_framework.md")); + assert.equal(manifest.research_profile, "budget-unfriendly"); + assert.equal(manifest.toolchain.quality_profile, "budget-unfriendly"); + assert.equal(manifest.toolchain.merge_strategy, "deep"); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("create celebrity accepts a string profile from runtime meta", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + const skillDir = createSkill( + baseDir, + "xu-zhisheng", + { + character: "celebrity", + name: "徐志胜", + display_name: "徐志胜", + classification: { language: "zh-CN" }, + profile: "中国脱口秀演员,以自嘲式观察喜剧著称。", + }, + "Work body", + "Persona body", + ); + + const savedMeta = readJson(join(skillDir, "meta.json")); + assert.equal(savedMeta.profile, "中国脱口秀演员,以自嘲式观察喜剧著称。"); + assert.match(savedMeta.summary, /中国脱口秀演员/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("existing dot-skill metadata keeps the legacy engine identifiers", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const skillDir = createSkill( + baseDir, + "legacy", + { + name: "Legacy", + preset: "dot.colleague.v1", + engine: { name: "dot-skill" }, + generation: { engine: "dot-skill" }, + artifacts: { + combined_name: "colleague_legacy", + work_name: "colleague_legacy_work", + persona_name: "colleague_legacy_persona", + }, + }, + "Work body", + "Persona body", + ); + + const savedMeta = readJson(join(skillDir, "meta.json")); + assert.equal(savedMeta.preset, "dot.colleague.v1"); + assert.equal(savedMeta.engine.name, "dot-skill"); + assert.equal(savedMeta.generation.engine, "dot-skill"); + assert.equal(savedMeta.artifacts.combined_name, "colleague_legacy"); + assert.match(readFileSync(join(skillDir, "SKILL.md"), "utf8"), /name: colleague_legacy/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update preserves names from legacy meta without artifacts", () => { + const root = tempDir(); + try { + const skillDir = join(root, "skills", "colleague", "legacy_person"); + mkdirSync(skillDir, { recursive: true }); + mkdirSync(join(skillDir, "versions")); + writeFileSync(join(skillDir, "meta.json"), JSON.stringify({ name: "Legacy Person", type: "colleague", version: "v1" }), "utf8"); + writeFileSync(join(skillDir, "work.md"), "Legacy work\n", "utf8"); + writeFileSync(join(skillDir, "persona.md"), "Legacy persona\n", "utf8"); + const legacyNames = { + "SKILL.md": "colleague_legacy_person", + "work_skill.md": "colleague_legacy_person_work", + "persona_skill.md": "colleague_legacy_person_persona", + }; + for (const [filename, name] of Object.entries(legacyNames)) { + writeFileSync(join(skillDir, filename), `---\nname: ${name}\ndescription: Legacy\n---\n\nLegacy body\n`, "utf8"); + } + + updateSkill(skillDir, "Updated work"); + + for (const [filename, name] of Object.entries(legacyNames)) { + assert.match(readFileSync(join(skillDir, filename), "utf8"), new RegExp(`name: ${name}`)); + } + assert.equal(readJson(join(skillDir, "meta.json")).artifacts.combined_command, "colleague-legacy-person"); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update rejects traversal in the stored version before backing up", () => { + const root = tempDir(); + try { + const skillDir = createSkill(join(root, "skills", "colleague"), "unsafe-version", { name: "Unsafe Version" }, "Work body", "Persona body"); + const metaPath = join(skillDir, "meta.json"); + const meta = readJson(metaPath); + meta.version = "../../../../escape"; + meta.lifecycle.version = "../../../../escape"; + writeFileSync(metaPath, JSON.stringify(meta), "utf8"); + + assert.throws(() => updateSkill(skillDir, "Should not be written"), /safe path segment/); + assert.equal(existsSync(join(root, "escape")), false); + assert.equal(readFileSync(join(skillDir, "work.md"), "utf8").includes("Should not be written"), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update regenerates the manifest and archives the artifacts", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const skillDir = createSkill(baseDir, "zhangsan", { name: "Eulalie" }, "Initial work", "Initial persona"); + + const newVersion = updateSkill(skillDir, "More work", null, { scene: "challenged", wrong: "apologize", correct: "ask for evidence" }); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const manifest = readJson(join(skillDir, "manifest.json")); + const personaDoc = readFileSync(join(skillDir, "persona.md"), "utf8"); + + assert.equal(newVersion, "v2"); + assert.equal(savedMeta.version, "v2"); + assert.equal(savedMeta.corrections_count, 1); + assert.ok(existsSync(join(skillDir, "versions", "v1", "manifest.json"))); + assert.equal(manifest.entrypoints.default, "SKILL.md"); + assert.match(personaDoc, /apologize/); + assert.match(personaDoc, /ask for evidence/); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update accepts multiple persona corrections in one payload", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + const skillDir = createSkill( + baseDir, + "zhou-qimo", + { character: "celebrity", name: "周奇墨", classification: { language: "zh-CN" } }, + "Initial work", + "Initial persona", + ); + + const newVersion = updateSkill(skillDir, null, null, { + persona_corrections: [ + { scene: "铺陈处境时", wrong: "一上来就下判断", correct: "先把处境讲得很普通,再轻轻点一下" }, + { scene: "表达立场时", wrong: "写成明显自嘲型", correct: "和观众一起承认大家都在局里" }, + ], + }); + + const savedMeta = readJson(join(skillDir, "meta.json")); + const personaDoc = readFileSync(join(skillDir, "persona.md"), "utf8"); + + assert.equal(newVersion, "v2"); + assert.equal(savedMeta.corrections_count, 2); + assert.match(personaDoc, /一上来就下判断/); + assert.match(personaDoc, /写成明显自嘲型/); + assert.equal(personaDoc.split("## Correction Log").length - 1, 1); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("update replaces existing markdown sections instead of appending duplicates", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "celebrity"); + const skillDir = createSkill( + baseDir, + "zhou-qimo", + { character: "celebrity", name: "周奇墨", classification: { language: "zh-CN" } }, + ["# Work", "", "## 表达规范", "", "- 原始表述", "", "## 输出风格", "", "- 原始结构"].join("\n"), + ["# Persona", "", "## Layer 2: Expression DNA", "", "旧内容", "", "## Layer 3: Mental Models", "", "保持不变"].join("\n"), + ); + + updateSkill( + skillDir, + ["## 表达规范", "", "- 新的节奏控制", "", "## 输出风格", "", "- 新的结构模板"].join("\n"), + ["## Layer 2: Expression DNA", "", "新内容"].join("\n"), + null, + ); + + const workDoc = readFileSync(join(skillDir, "work.md"), "utf8"); + const personaDoc = readFileSync(join(skillDir, "persona.md"), "utf8"); + + assert.equal(workDoc.split("## 表达规范").length - 1, 1); + assert.equal(workDoc.split("## 输出风格").length - 1, 1); + assert.match(workDoc, /新的节奏控制/); + assert.equal(workDoc.includes("原始表述"), false); + assert.equal(personaDoc.split("## Layer 2: Expression DNA").length - 1, 1); + assert.match(personaDoc, /新内容/); + assert.equal(personaDoc.includes("旧内容"), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("backup and rollback include the manifest", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + const skillDir = createSkill(baseDir, "zhangsan", { name: "Eulalie" }, "v1 work", "v1 persona"); + + backupCurrentVersion(skillDir); + updateSkill(skillDir, "v2 work"); + + const success = rollback(skillDir, "v1"); + const restoredWork = readFileSync(join(skillDir, "work.md"), "utf8"); + + assert.equal(success, true); + assert.match(restoredWork, /v1 work/); + assert.ok(existsSync(join(skillDir, "versions", "v1", "manifest.json"))); + assert.equal(rollback(skillDir, "../v1"), false); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("the version manager still resolves the legacy colleagues root", () => { + const root = tempDir(); + const cwd = process.cwd(); + try { + process.chdir(root); + createSkill("colleagues", "zhangsan", { name: "Eulalie" }, "v1 work", "v1 persona"); + assert.equal(resolveExistingStorageRoot("colleague", "zhangsan"), "colleagues"); + assert.deepEqual(listSkills("colleagues").map((skill) => skill.slug), ["zhangsan"]); + } finally { + process.chdir(cwd); + rmSync(root, { recursive: true, force: true }); + } +}); + +test("character prompt bundles exist", () => { + for (const character of ["colleague", "relationship", "celebrity"]) { + const preset = getCharacterPreset(character); + for (const promptPath of Object.values(preset.prompt_bundle)) { + if (typeof promptPath !== "string" || !promptPath.startsWith("prompts/")) continue; + assert.ok(existsSync(join(projectRoot, promptPath)), `missing prompt file for ${character}: ${promptPath}`); + } + // v2: research tools are CLI commands; anything that is still a path must + // exist, and every command must be registered (or at least planned). + const knownCommands = new Set([...listCommands(), ...Object.keys(PLANNED)]); + for (const [tool, value] of Object.entries(preset.research_tools ?? {})) { + if (typeof value !== "string") continue; + if (!value.startsWith("distilly ")) { + assert.ok(existsSync(join(projectRoot, value)), `missing research tool for ${character}: ${value}`); + continue; + } + const [first, second] = value.slice("distilly ".length).split(" "); + const resolved = knownCommands.has(`${first} ${second}`) ? `${first} ${second}` : first; + assert.ok( + knownCommands.has(resolved), + `${character}.research_tools.${tool} names an unregistered command: ${value}`, + ); + } + for (const profileName of Object.keys(preset.research_profiles ?? {})) { + const profile = getResearchProfilePreset(character, profileName); + for (const promptPath of Object.values(profile.prompt_bundle ?? {})) { + if (typeof promptPath !== "string" || !promptPath.startsWith("prompts/")) continue; + assert.ok(existsSync(join(projectRoot, promptPath)), `missing profile prompt for ${character}/${profileName}: ${promptPath}`); + } + for (const referencePath of profile.references ?? []) { + assert.ok(existsSync(join(projectRoot, referencePath)), `missing profile reference for ${character}/${profileName}: ${referencePath}`); + } + } + } + + const friendlyPrompt = readFileSync(join(projectRoot, "prompts", "celebrity", "research.md"), "utf8"); + assert.match(friendlyPrompt, /01_core_profile\.md/); + assert.match(friendlyPrompt, /03_expression_and_reception\.md/); + assert.match(friendlyPrompt.toLowerCase(), /do not collapse the whole pass into one monolithic note/); + assert.match(friendlyPrompt, /actual inspected pages/); + assert.match(friendlyPrompt, /tools\/research\/xquik_public_posts\.py/); + assert.match(friendlyPrompt.split(/\s+/).join(" "), /untrusted candidate evidence/); + + const strictPrompt = readFileSync(join(projectRoot, "prompts", "celebrity", "budget_unfriendly", "research.md"), "utf8"); + assert.match(strictPrompt, /01_writings\.md/); + assert.match(strictPrompt, /06_timeline\.md/); + assert.match(strictPrompt, /at least 8 grounded source URLs/); + assert.match(strictPrompt, /Do not replace these six files with one merged scratchpad/); + assert.match(strictPrompt, /actual inspected pages/); + assert.match(strictPrompt, /tools\/research\/xquik_public_posts\.py/); + assert.match(strictPrompt.split(/\s+/).join(" "), /untrusted candidate evidence/); +}); + +test("a renamed skill directory stays addressable through renameSync", () => { + const root = tempDir(); + try { + const baseDir = join(root, "skills", "colleague"); + createSkill(baseDir, "legacy", { name: "Zadie Smith" }, "Work body", "Persona body"); + renameSync(join(baseDir, "legacy"), join(baseDir, "Zadie Smith")); + const skills = listSkills(baseDir); + assert.equal(skills.length, 1); + assert.equal(skills[0].slug, "Zadie Smith"); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); From b13d5bfc0c5622a564d43cad4032dd15404676eb Mon Sep 17 00:00:00 2001 From: zhoutianyi Date: Tue, 15 Sep 2026 15:37:59 +0800 Subject: [PATCH 14/15] =?UTF-8?q?chore(python):=20=E9=80=80=E5=BD=B9=2037?= =?UTF-8?q?=20=E4=B8=AA=20Python=20=E6=96=87=E4=BB=B6?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 重建说明:原分支删除这批 Python 文件是分几个提交做的,这里合并为一个删除提交;文件清单来自 git diff 的 D 状态记录。 --- requirements.txt | 15 - tests/test_cli_lifecycle.py | 609 ------------ tests/test_config_migration.py | 53 - tests/test_install_claude_generated_skill.py | 98 -- tests/test_install_generated_skill.py | 126 --- tests/test_install_hermes_skill.py | 83 -- tests/test_install_openclaw_and_codex.py | 140 --- tests/test_research_tools.py | 278 ------ tests/test_skill_entrypoint_docs.py | 328 ------- tests/test_skill_writer.py | 719 -------------- tests/test_xquik_public_posts.py | 242 ----- tools/dingtalk_auto_collector.py | 790 --------------- tools/email_parser.py | 339 ------- tools/feishu_auto_collector.py | 960 ------------------- tools/feishu_browser.py | 374 -------- tools/feishu_mcp_client.py | 314 ------ tools/feishu_parser.py | 251 ----- tools/install_claude_generated_skill.py | 108 --- tools/install_codex_generated_skill.py | 57 -- tools/install_codex_skill.py | 65 -- tools/install_generated_skill.py | 76 -- tools/install_generated_skill_common.py | 130 --- tools/install_hermes_skill.py | 65 -- tools/install_openclaw_generated_skill.py | 57 -- tools/install_openclaw_skill.py | 65 -- tools/research/__init__.py | 1 - tools/research/download_subtitles.sh | 31 - tools/research/merge_research.py | 277 ------ tools/research/quality_check.py | 253 ----- tools/research/srt_to_transcript.py | 92 -- tools/research/transcribe_audio.py | 304 ------ tools/research/xquik_public_posts.py | 406 -------- tools/skill_presets.py | 288 ------ tools/skill_schema.py | 411 -------- tools/skill_writer.py | 705 -------------- tools/slack_auto_collector.py | 722 -------------- tools/version_manager.py | 239 ----- 37 files changed, 10071 deletions(-) delete mode 100644 requirements.txt delete mode 100644 tests/test_cli_lifecycle.py delete mode 100644 tests/test_config_migration.py delete mode 100644 tests/test_install_claude_generated_skill.py delete mode 100644 tests/test_install_generated_skill.py delete mode 100644 tests/test_install_hermes_skill.py delete mode 100644 tests/test_install_openclaw_and_codex.py delete mode 100644 tests/test_research_tools.py delete mode 100644 tests/test_skill_entrypoint_docs.py delete mode 100644 tests/test_skill_writer.py delete mode 100644 tests/test_xquik_public_posts.py delete mode 100644 tools/dingtalk_auto_collector.py delete mode 100644 tools/email_parser.py delete mode 100644 tools/feishu_auto_collector.py delete mode 100644 tools/feishu_browser.py delete mode 100644 tools/feishu_mcp_client.py delete mode 100644 tools/feishu_parser.py delete mode 100644 tools/install_claude_generated_skill.py delete mode 100644 tools/install_codex_generated_skill.py delete mode 100644 tools/install_codex_skill.py delete mode 100755 tools/install_generated_skill.py delete mode 100644 tools/install_generated_skill_common.py delete mode 100755 tools/install_hermes_skill.py delete mode 100644 tools/install_openclaw_generated_skill.py delete mode 100644 tools/install_openclaw_skill.py delete mode 100644 tools/research/__init__.py delete mode 100755 tools/research/download_subtitles.sh delete mode 100755 tools/research/merge_research.py delete mode 100755 tools/research/quality_check.py delete mode 100755 tools/research/srt_to_transcript.py delete mode 100755 tools/research/transcribe_audio.py delete mode 100755 tools/research/xquik_public_posts.py delete mode 100644 tools/skill_presets.py delete mode 100644 tools/skill_schema.py delete mode 100644 tools/skill_writer.py delete mode 100644 tools/slack_auto_collector.py delete mode 100644 tools/version_manager.py diff --git a/requirements.txt b/requirements.txt deleted file mode 100644 index 19dea998..00000000 --- a/requirements.txt +++ /dev/null @@ -1,15 +0,0 @@ -# Required -requests>=2.28.0 - -# Optional: Chinese name → slug conversion -pypinyin>=0.48.0 - -# Optional: Playwright for Lark-compatible browser login / DingTalk message scraping -playwright>=1.40.0 - -# Optional: Slack auto collector -slack-sdk>=3.27.0 - -# Optional: Word/Excel parsing (convert to PDF/CSV first if unavailable) -python-docx>=1.1.0 -openpyxl>=3.1.0 diff --git a/tests/test_cli_lifecycle.py b/tests/test_cli_lifecycle.py deleted file mode 100644 index 65065090..00000000 --- a/tests/test_cli_lifecycle.py +++ /dev/null @@ -1,609 +0,0 @@ -from __future__ import annotations - -import json -import os -import subprocess -import sys -import tempfile -import unittest -from pathlib import Path - - -PROJECT_ROOT = Path(__file__).resolve().parents[1] -PYTHON = sys.executable - - -class CliLifecycleTest(unittest.TestCase): - def run_cmd( - self, - *args: str, - cwd: Path | None = None, - env: dict[str, str] | None = None, - ) -> subprocess.CompletedProcess[str]: - merged_env = os.environ.copy() - merged_env.setdefault("DISTILLY_AUTO_INSTALL_CLAUDE", "0") - if env: - merged_env.update(env) - return subprocess.run( - list(args), - cwd=cwd or PROJECT_ROOT, - text=True, - capture_output=True, - check=True, - env=merged_env, - ) - - def write_json(self, path: Path, payload: dict) -> Path: - path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8") - return path - - def test_default_colleague_cli_uses_skills_colleague_root(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - work_path = tmp_root / "work.md" - persona_path = tmp_root / "persona.md" - meta_path = tmp_root / "meta.json" - - work_path.write_text("Work body\n", encoding="utf-8") - persona_path.write_text("Persona body\n", encoding="utf-8") - self.write_json( - meta_path, - { - "character": "colleague", - "display_name": "Eulalie", - "classification": {"language": "en"}, - }, - ) - - create = self.run_cmd( - PYTHON, - str(PROJECT_ROOT / "tools" / "skill_writer.py"), - "--action", - "create", - "--character", - "colleague", - "--slug", - "eulalie", - "--name", - "Eulalie", - "--meta", - str(meta_path), - "--work", - str(work_path), - "--persona", - str(persona_path), - cwd=tmp_root, - ) - - self.assertIn("Created skill:", create.stdout) - self.assertTrue((tmp_root / "skills" / "colleague" / "eulalie" / "SKILL.md").exists()) - - def test_claude_auto_install_is_opt_in_with_legacy_env_compatibility(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - home = root / "home" - base_dir = root / "skills" / "colleague" - home.mkdir() - base_dir.mkdir(parents=True) - meta_path = self.write_json( - root / "meta.json", - { - "character": "colleague", - "display_name": "Eulalie", - "classification": {"language": "en"}, - }, - ) - work_path = root / "work.md" - persona_path = root / "persona.md" - work_path.write_text("Work body\n", encoding="utf-8") - persona_path.write_text("Persona body\n", encoding="utf-8") - - def create_with_env( - slug: str, - settings: dict[str, str], - *extra_args: str, - ) -> Path: - env = os.environ.copy() - env["HOME"] = str(home) - env.pop("DISTILLY_AUTO_INSTALL_CLAUDE", None) - env.pop("DOT_SKILL_AUTO_INSTALL_CLAUDE", None) - env.update(settings) - subprocess.run( - [ - PYTHON, - str(PROJECT_ROOT / "tools" / "skill_writer.py"), - "--action", - "create", - "--character", - "colleague", - "--slug", - slug, - "--name", - "Eulalie", - "--meta", - str(meta_path), - "--work", - str(work_path), - "--persona", - str(persona_path), - "--base-dir", - str(base_dir), - *extra_args, - ], - cwd=PROJECT_ROOT, - text=True, - capture_output=True, - check=True, - env=env, - ) - return home / ".claude" / "skills" / f"colleague-{slug}" / "SKILL.md" - - self.assertFalse(create_with_env("default-off", {}).exists()) - self.assertTrue( - create_with_env( - "legacy-on", - {"DOT_SKILL_AUTO_INSTALL_CLAUDE": "1"}, - ).exists() - ) - self.assertFalse( - create_with_env( - "legacy-off", - {"DOT_SKILL_AUTO_INSTALL_CLAUDE": "0"}, - ).exists() - ) - self.assertFalse( - create_with_env( - "new-wins-off", - { - "DISTILLY_AUTO_INSTALL_CLAUDE": "0", - "DOT_SKILL_AUTO_INSTALL_CLAUDE": "1", - }, - ).exists() - ) - self.assertTrue( - create_with_env( - "new-wins-on", - { - "DISTILLY_AUTO_INSTALL_CLAUDE": "1", - "DOT_SKILL_AUTO_INSTALL_CLAUDE": "0", - }, - ).exists() - ) - self.assertFalse( - create_with_env( - "explicit-off", - {"DISTILLY_AUTO_INSTALL_CLAUDE": "1"}, - "--no-install-claude-skill", - ).exists() - ) - - def test_create_name_only_normalizes_slug_and_rejects_unsafe_explicit_slug(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - writer = str(PROJECT_ROOT / "tools" / "skill_writer.py") - - self.run_cmd( - PYTHON, - writer, - "--action", - "create", - "--name", - "Zadie Smith", - "--base-dir", - "skills/colleague", - cwd=root, - ) - generated = root / "skills" / "colleague" / "zadie-smith" / "SKILL.md" - self.assertIn("name: colleague-zadie-smith", generated.read_text(encoding="utf-8")) - - with self.assertRaises(subprocess.CalledProcessError): - self.run_cmd( - PYTHON, - writer, - "--action", - "create", - "--slug", - "../escape", - "--base-dir", - "skills/colleague", - cwd=root, - ) - self.assertFalse((root / "skills" / "escape").exists()) - - def test_update_accepts_safe_legacy_slug_with_spaces(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - writer = str(PROJECT_ROOT / "tools" / "skill_writer.py") - base_dir = root / "skills" / "colleague" - - self.run_cmd( - PYTHON, - writer, - "--action", - "create", - "--slug", - "legacy", - "--name", - "Zadie Smith", - "--base-dir", - str(base_dir), - cwd=root, - ) - legacy_dir = base_dir / "Zadie Smith" - (base_dir / "legacy").rename(legacy_dir) - work_patch = root / "work-patch.md" - work_patch.write_text("## Update\n\nLegacy directory remains addressable.\n", encoding="utf-8") - - update = self.run_cmd( - PYTHON, - writer, - "--action", - "update", - "--slug", - "Zadie Smith", - "--base-dir", - str(base_dir), - "--work-patch", - str(work_patch), - cwd=root, - ) - - self.assertIn("Updated skill to v2:", update.stdout) - self.assertIn("Zadie Smith", update.stdout) - self.assertIn("Legacy directory remains addressable", (legacy_dir / "work.md").read_text(encoding="utf-8")) - saved_meta = json.loads( - (legacy_dir / "meta.json").read_text(encoding="utf-8") - ) - self.assertEqual( - saved_meta["artifacts"]["combined_command"], - "colleague-zadie-smith", - ) - - def test_version_manager_rejects_slug_and_version_traversal(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - base_dir = root / "skills" / "colleague" - victim_versions = root / "skills" / "victim" / "versions" - for index in range(11): - (victim_versions / f"v{index}").mkdir(parents=True) - manager = str(PROJECT_ROOT / "tools" / "version_manager.py") - - with self.assertRaises(subprocess.CalledProcessError): - self.run_cmd( - PYTHON, - manager, - "--action", - "cleanup", - "--slug", - "../victim", - "--base-dir", - str(base_dir), - cwd=root, - ) - self.assertEqual(len(list(victim_versions.iterdir())), 11) - - self.run_cmd( - PYTHON, - str(PROJECT_ROOT / "tools" / "skill_writer.py"), - "--action", - "create", - "--slug", - "safe", - "--name", - "Safe", - "--base-dir", - str(base_dir), - cwd=root, - ) - with self.assertRaises(subprocess.CalledProcessError): - self.run_cmd( - PYTHON, - manager, - "--action", - "rollback", - "--slug", - "safe", - "--version", - "../victim", - "--base-dir", - str(base_dir), - cwd=root, - ) - self.assertTrue((base_dir / "safe" / "SKILL.md").exists()) - - def test_character_lifecycle_via_cli(self) -> None: - fixtures = { - "colleague": { - "name": "Eulalie", - "slug": "eulalie", - "base_dir": "skills/colleague", - }, - "relationship": { - "name": "Mireille", - "slug": "mireille", - "base_dir": "skills/relationship", - }, - "celebrity": { - "name": "Zadie Smith", - "slug": "zadie-smith", - "base_dir": "skills/celebrity", - }, - } - - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - - for character, fixture in fixtures.items(): - base_dir = root / fixture["base_dir"] - base_dir.mkdir(parents=True, exist_ok=True) - meta_path = root / f"{fixture['slug']}_meta.json" - work_path = root / f"{fixture['slug']}_work.md" - persona_path = root / f"{fixture['slug']}_persona.md" - work_patch_path = root / f"{fixture['slug']}_work_patch.md" - correction_path = root / f"{fixture['slug']}_correction.json" - - self.write_json( - meta_path, - { - "character": character, - "display_name": fixture["name"], - "classification": {"language": "en"}, - "profile": {"role": "Builder"}, - "tags": {"personality": ["precise", "skeptical"]}, - "knowledge_sources": ["manual-notes"], - }, - ) - work_path.write_text( - "\n".join( - [ - "## mental models", - "- First-principles reasoning", - "- Skeptical framing", - "- Long-horizon tradeoffs", - "", - "## limitations", - "- Avoids operational detail", - "", - "Sources:", - "https://example.com/articles/long-form-profile", - "https://example.com/interviews/episode-42", - ] - ) - + "\n", - encoding="utf-8", - ) - persona_path.write_text( - "\n".join( - [ - "## expression DNA", - "- Sentence rhythm is clipped.", - "- Uses metaphor when disagreeing.", - "", - "## honest boundaries", - "- States what they do not know.", - "", - "## contradictions", - "- Alternates between certainty and doubt.", - ] - ) - + "\n", - encoding="utf-8", - ) - work_patch_path.write_text("## new evidence\n- Adds a later example.\n", encoding="utf-8") - self.write_json( - correction_path, - { - "scene": "disagreement", - "wrong": "flatten disagreement into politeness", - "correct": "surface the disagreement and justify it directly", - }, - ) - - create = self.run_cmd( - PYTHON, - "tools/skill_writer.py", - "--action", - "create", - "--character", - character, - "--slug", - fixture["slug"], - "--name", - fixture["name"], - "--meta", - str(meta_path), - "--work", - str(work_path), - "--persona", - str(persona_path), - "--base-dir", - str(base_dir), - ) - self.assertIn("Created skill:", create.stdout) - - skill_dir = base_dir / fixture["slug"] - self.assertTrue((skill_dir / "SKILL.md").exists()) - self.assertTrue((skill_dir / "manifest.json").exists()) - - list_result = self.run_cmd( - PYTHON, - "tools/skill_writer.py", - "--action", - "list", - "--character", - character, - "--base-dir", - str(base_dir), - ) - self.assertIn(fixture["slug"], list_result.stdout) - self.assertIn(f"Character: {character}", list_result.stdout) - - update = self.run_cmd( - PYTHON, - "tools/skill_writer.py", - "--action", - "update", - "--character", - character, - "--slug", - fixture["slug"], - "--work-patch", - str(work_patch_path), - "--correction-json", - str(correction_path), - "--base-dir", - str(base_dir), - ) - self.assertIn("Updated skill to v2", update.stdout) - - versions = self.run_cmd( - PYTHON, - "tools/version_manager.py", - "--action", - "list", - "--character", - character, - "--slug", - fixture["slug"], - "--base-dir", - str(base_dir), - ) - self.assertIn("v1", versions.stdout) - - rollback = self.run_cmd( - PYTHON, - "tools/version_manager.py", - "--action", - "rollback", - "--character", - character, - "--slug", - fixture["slug"], - "--version", - "v1", - "--base-dir", - str(base_dir), - ) - self.assertIn("rolled back to v1", rollback.stdout) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - self.assertEqual(saved_meta["character"], character) - self.assertTrue(saved_meta["version"].startswith("v1")) - - combined_skill = (skill_dir / "SKILL.md").read_text(encoding="utf-8") - self.assertIn("## PART A: Work", combined_skill) - self.assertIn("## PART B: Persona", combined_skill) - - if character == "celebrity": - subtitle_path = skill_dir / "knowledge" / "subtitles" / "sample.vtt" - subtitle_path.write_text( - "WEBVTT\n\n00:00:00.000 --> 00:00:01.000\nHello\n\n" - "00:00:01.000 --> 00:00:02.000\nworld.\n", - encoding="utf-8", - ) - transcript_path = skill_dir / "knowledge" / "transcripts" / "sample.txt" - transcript = self.run_cmd( - PYTHON, - "tools/research/srt_to_transcript.py", - str(subtitle_path), - str(transcript_path), - ) - self.assertIn(str(transcript_path), transcript.stdout) - self.assertIn("Hello world.", transcript_path.read_text(encoding="utf-8")) - - raw_note = skill_dir / "knowledge" / "research" / "raw" / "01.md" - raw_note.write_text( - "# Notes\n" - "- Strong focus on first-person essays\n" - "- Repeats a skeptical framing\n" - "https://example.com/essays/first-person-observation\n" - "primary source\n", - encoding="utf-8", - ) - merged = self.run_cmd( - PYTHON, - "tools/research/merge_research.py", - str(skill_dir), - ) - self.assertIn("summary.md", merged.stdout) - summary_text = (skill_dir / "knowledge" / "research" / "merged" / "summary.md").read_text( - encoding="utf-8" - ) - self.assertIn("Research Summary", summary_text) - - quality = self.run_cmd( - PYTHON, - "tools/research/quality_check.py", - str(skill_dir / "SKILL.md"), - ) - self.assertIn("OVERALL PASS", quality.stdout) - - def test_cli_can_install_generated_skill_into_supported_host_paths(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - base_dir = root / "skills" / "celebrity" - base_dir.mkdir(parents=True, exist_ok=True) - - meta_path = root / "zhou_qimo_meta.json" - work_path = root / "zhou_qimo_work.md" - persona_path = root / "zhou_qimo_persona.md" - self.write_json( - meta_path, - { - "character": "celebrity", - "display_name": "周奇墨", - "classification": {"language": "zh-CN"}, - }, - ) - work_path.write_text("Work body\n", encoding="utf-8") - persona_path.write_text("Persona body\n", encoding="utf-8") - - claude_skills_dir = root / ".claude" / "skills" - claude_commands_dir = root / ".claude" / "commands" - openclaw_skills_dir = root / ".openclaw" / "workspace" / "skills" - codex_skills_dir = root / ".agents" / "skills" - - create = self.run_cmd( - PYTHON, - "tools/skill_writer.py", - "--action", - "create", - "--character", - "celebrity", - "--slug", - "zhou-qimo", - "--name", - "周奇墨", - "--meta", - str(meta_path), - "--work", - str(work_path), - "--persona", - str(persona_path), - "--base-dir", - str(base_dir), - "--install-claude-skill", - "--install-claude-command-shim", - "--claude-skills-dir", - str(claude_skills_dir), - "--claude-commands-dir", - str(claude_commands_dir), - "--install-openclaw-skill", - "--openclaw-skills-dir", - str(openclaw_skills_dir), - "--install-codex-skill", - "--codex-skills-dir", - str(codex_skills_dir), - ) - - self.assertIn("Claude trigger: /celebrity-zhou-qimo", create.stdout) - self.assertIn("OpenClaw trigger: /celebrity-zhou-qimo", create.stdout) - self.assertIn("Codex skill name: celebrity-zhou-qimo", create.stdout) - self.assertTrue((claude_skills_dir / "celebrity-zhou-qimo" / "SKILL.md").exists()) - self.assertTrue((claude_commands_dir / "celebrity-zhou-qimo.md").exists()) - self.assertTrue((openclaw_skills_dir / "celebrity-zhou-qimo" / "SKILL.md").exists()) - self.assertTrue((codex_skills_dir / "celebrity-zhou-qimo" / "SKILL.md").exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_config_migration.py b/tests/test_config_migration.py deleted file mode 100644 index 5f19b066..00000000 --- a/tests/test_config_migration.py +++ /dev/null @@ -1,53 +0,0 @@ -from __future__ import annotations - -import json -import stat -import sys -import tempfile -import unittest -from pathlib import Path -from unittest.mock import patch - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -import dingtalk_auto_collector # noqa: E402 -import feishu_auto_collector # noqa: E402 -import feishu_mcp_client # noqa: E402 - - -COLLECTOR_MODULES = ( - dingtalk_auto_collector, - feishu_auto_collector, - feishu_mcp_client, -) - - -class ConfigMigrationTest(unittest.TestCase): - def test_collectors_read_legacy_config_and_save_new_config_privately(self) -> None: - for module in COLLECTOR_MODULES: - with self.subTest(module=module.__name__), tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - current = root / ".distilly" / "config.json" - legacy = root / ".colleague-skill" / "config.json" - legacy.parent.mkdir(parents=True) - legacy.write_text(json.dumps({"source": "legacy"}), encoding="utf-8") - - with ( - patch.object(module, "CONFIG_PATH", current), - patch.object(module, "LEGACY_CONFIG_PATH", legacy), - ): - self.assertEqual(module.load_config(), {"source": "legacy"}) - module.save_config({"source": "distilly"}) - - self.assertEqual( - json.loads(current.read_text(encoding="utf-8")), - {"source": "distilly"}, - ) - self.assertEqual(stat.S_IMODE(current.stat().st_mode), 0o600) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_install_claude_generated_skill.py b/tests/test_install_claude_generated_skill.py deleted file mode 100644 index 25dac1e6..00000000 --- a/tests/test_install_claude_generated_skill.py +++ /dev/null @@ -1,98 +0,0 @@ -from __future__ import annotations - -import tempfile -import unittest -from pathlib import Path -import sys - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from install_claude_generated_skill import ( # noqa: E402 - install_generated_skill, - should_install_command_shim, -) -import skill_writer # noqa: E402 - - -class ClaudeGeneratedSkillInstallTest(unittest.TestCase): - def test_install_generated_skill_writes_claude_skill_folder(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - generated_root = tmp_root / "skills" / "celebrity" - claude_skills = tmp_root / ".claude" / "skills" - - skill_dir = skill_writer.create_skill( - generated_root, - "zhou-qimo", - { - "character": "celebrity", - "name": "周奇墨", - "classification": {"language": "zh-CN"}, - }, - "Work body", - "Persona body", - ) - - result = install_generated_skill( - skill_dir, - claude_skills, - force=True, - ) - - installed_file = claude_skills / "celebrity-zhou-qimo" / "SKILL.md" - metadata_file = claude_skills / "celebrity-zhou-qimo" / ".distilly-install.json" - - self.assertEqual(result["command_name"], "celebrity-zhou-qimo") - self.assertTrue(installed_file.exists()) - self.assertTrue(metadata_file.exists()) - self.assertIn( - "name: celebrity-zhou-qimo", - installed_file.read_text(encoding="utf-8"), - ) - self.assertFalse(result["command_shim_installed"]) - - def test_install_generated_skill_can_write_windows_command_shim(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - generated_root = tmp_root / "skills" / "relationship" - claude_skills = tmp_root / ".claude" / "skills" - claude_commands = tmp_root / ".claude" / "commands" - - skill_dir = skill_writer.create_skill( - generated_root, - "mireille", - { - "character": "relationship", - "name": "Mireille", - }, - "Work body", - "Persona body", - ) - - result = install_generated_skill( - skill_dir, - claude_skills, - commands_dir=claude_commands, - force=True, - install_command_shim=True, - ) - - command_file = claude_commands / "relationship-mireille.md" - self.assertTrue(result["command_shim_installed"]) - self.assertEqual(result["command_path"], command_file) - self.assertTrue(command_file.exists()) - self.assertIn( - "name: relationship-mireille", - command_file.read_text(encoding="utf-8"), - ) - - def test_windows_detection_only_enables_command_shim_on_windows(self) -> None: - self.assertTrue(should_install_command_shim("Windows")) - self.assertFalse(should_install_command_shim("Darwin")) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_install_generated_skill.py b/tests/test_install_generated_skill.py deleted file mode 100644 index f2535ad2..00000000 --- a/tests/test_install_generated_skill.py +++ /dev/null @@ -1,126 +0,0 @@ -from __future__ import annotations - -import json -import tempfile -import unittest -from pathlib import Path -import sys - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from install_generated_skill import default_skills_dir # noqa: E402 -from install_generated_skill_common import install_generated_skill # noqa: E402 -import skill_writer # noqa: E402 - - -class GeneratedSkillInstallTest(unittest.TestCase): - def create_legacy_skill(self, root: Path) -> Path: - skill_dir = skill_writer.create_skill( - root, - "mireille", - {"character": "relationship", "name": "Mireille"}, - "Work body", - "Persona body", - ) - skill_path = skill_dir / "SKILL.md" - skill_path.write_text( - skill_path.read_text(encoding="utf-8").replace( - "name: relationship-mireille", - "name: relationship_mireille", - 1, - ), - encoding="utf-8", - ) - meta_path = skill_dir / "meta.json" - meta = json.loads(meta_path.read_text(encoding="utf-8")) - meta.pop("artifacts", None) - meta_path.write_text(json.dumps(meta), encoding="utf-8") - return skill_dir - - def test_default_skills_dirs_cover_all_documented_hosts(self) -> None: - home = Path("/example/home") - expected = { - "claude-code": home / ".claude" / "skills", - "openclaw": home / ".openclaw" / "workspace" / "skills", - "hermes": home / ".hermes" / "skills" / "distilly-generated", - "codex": home / ".agents" / "skills", - "deepseek-harness": home / ".dsh" / "skills", - "pi": home / ".pi" / "agent" / "skills", - "grok-build": home / ".grok" / "skills", - "opencode": home / ".config" / "opencode" / "skills", - } - self.assertEqual( - {host: default_skills_dir(host, home, {}) for host in expected}, - expected, - ) - self.assertEqual( - default_skills_dir( - "deepseek-harness", - home, - {"DSH_HOME": "/custom/dsh"}, - ), - Path("/custom/dsh/skills"), - ) - - def test_install_rewrites_only_the_legacy_copy_to_canonical_name(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - source = self.create_legacy_skill(root / "generated") - skills_dir = root / ".hermes" / "skills" / "distilly-generated" - - result = install_generated_skill( - source, - skills_dir, - force=True, - host="hermes", - ) - - installed = result["skill_dir"] - self.assertEqual(installed.name, "relationship-mireille") - self.assertEqual( - {path.name for path in installed.iterdir()}, - {"SKILL.md", ".distilly-install.json"}, - ) - self.assertIn( - "name: relationship-mireille", - (installed / "SKILL.md").read_text(encoding="utf-8"), - ) - self.assertIn( - "name: relationship_mireille", - (source / "SKILL.md").read_text(encoding="utf-8"), - ) - - def test_install_rejects_ancestor_or_descendant_destinations(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - source = self.create_legacy_skill(root / "generated") - - with self.assertRaisesRegex(ValueError, "must not overlap"): - install_generated_skill( - source, - source, - force=True, - host="test", - ) - - ancestor_install = root / "host" / "relationship-bundle" - ancestor_install.mkdir(parents=True) - nested_source = ancestor_install / "bundle" - source.rename(nested_source) - with self.assertRaisesRegex(ValueError, "must not overlap"): - install_generated_skill( - nested_source, - ancestor_install.parent, - force=True, - host="test", - ) - - self.assertTrue((nested_source / "meta.json").exists()) - self.assertTrue((nested_source / "work.md").exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_install_hermes_skill.py b/tests/test_install_hermes_skill.py deleted file mode 100644 index d4dabc1b..00000000 --- a/tests/test_install_hermes_skill.py +++ /dev/null @@ -1,83 +0,0 @@ -from __future__ import annotations - -import tempfile -import unittest -from pathlib import Path -import sys - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from install_hermes_skill import install_skill # noqa: E402 - - -class HermesInstallTest(unittest.TestCase): - def test_install_skill_copies_repo_layout(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - source = Path(tmp_dir) / "source" - destination = Path(tmp_dir) / "dest" / "distilly" - source.mkdir() - (source / "SKILL.md").write_text("name: distilly\n", encoding="utf-8") - (source / "README.md").write_text("# Distilly\n", encoding="utf-8") - - installed = install_skill(source, destination) - self.assertEqual(installed, destination) - self.assertTrue((destination / "SKILL.md").exists()) - self.assertTrue((destination / "README.md").exists()) - - def test_install_skill_dry_run_does_not_write(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - source = Path(tmp_dir) / "source" - destination = Path(tmp_dir) / "dest" / "distilly" - source.mkdir() - (source / "SKILL.md").write_text("name: distilly\n", encoding="utf-8") - - install_skill(source, destination, dry_run=True) - self.assertFalse(destination.exists()) - - def test_install_skill_dry_run_allows_existing_destination(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - source = Path(tmp_dir) / "source" - destination = Path(tmp_dir) / "dest" / "distilly" - source.mkdir(parents=True) - destination.mkdir(parents=True) - (source / "SKILL.md").write_text("name: distilly\n", encoding="utf-8") - - result = install_skill(source, destination, dry_run=True) - self.assertEqual(result, destination) - self.assertTrue(destination.exists()) - - def test_install_skill_does_not_delete_source_when_it_is_already_the_destination(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - source = Path(tmp_dir) / "distilly" - source.mkdir() - skill_file = source / "SKILL.md" - skill_file.write_text("name: distilly\n", encoding="utf-8") - - result = install_skill(source, source, force=True) - - self.assertEqual(result, source) - self.assertTrue(skill_file.exists()) - - def test_install_skill_rejects_nested_or_ancestor_destination(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - source = root / "parent" / "source" - source.mkdir(parents=True) - skill_file = source / "SKILL.md" - skill_file.write_text("name: distilly\n", encoding="utf-8") - nested = source / ".hermes" / "skills" / "distilly" - - with self.assertRaisesRegex(ValueError, "must not overlap"): - install_skill(source, nested, force=True) - with self.assertRaisesRegex(ValueError, "must not overlap"): - install_skill(source, source.parent, force=True) - - self.assertTrue(skill_file.exists()) - self.assertFalse(nested.exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_install_openclaw_and_codex.py b/tests/test_install_openclaw_and_codex.py deleted file mode 100644 index 2e41b39f..00000000 --- a/tests/test_install_openclaw_and_codex.py +++ /dev/null @@ -1,140 +0,0 @@ -from __future__ import annotations - -import tempfile -import unittest -from pathlib import Path -import sys - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from install_codex_generated_skill import install_generated_skill as install_codex_generated_skill # noqa: E402 -from install_codex_skill import install_skill as install_codex_skill # noqa: E402 -from install_openclaw_generated_skill import install_generated_skill as install_openclaw_generated_skill # noqa: E402 -from install_openclaw_skill import install_skill as install_openclaw_skill # noqa: E402 -import skill_writer # noqa: E402 - - -class OpenClawAndCodexInstallTest(unittest.TestCase): - def test_openclaw_and_codex_repo_installers_copy_repo_layout(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - source = tmp_root / "source" - source.mkdir() - (source / "SKILL.md").write_text("name: distilly\n", encoding="utf-8") - (source / "README.md").write_text("# Distilly\n", encoding="utf-8") - - openclaw_dest = tmp_root / "openclaw" / "distilly" - codex_dest = tmp_root / "agents" / "distilly" - - installed_openclaw = install_openclaw_skill(source, openclaw_dest) - installed_codex = install_codex_skill(source, codex_dest) - - self.assertEqual(installed_openclaw, openclaw_dest) - self.assertEqual(installed_codex, codex_dest) - self.assertTrue((openclaw_dest / "SKILL.md").exists()) - self.assertTrue((codex_dest / "SKILL.md").exists()) - - def test_repo_installers_do_not_delete_source_when_it_is_already_the_destination(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - source = Path(tmp_dir) / "distilly" - source.mkdir() - skill_file = source / "SKILL.md" - skill_file.write_text("name: distilly\n", encoding="utf-8") - - self.assertEqual(install_openclaw_skill(source, source, force=True), source) - self.assertEqual(install_codex_skill(source, source, force=True), source) - self.assertTrue(skill_file.exists()) - - def test_repo_installers_reject_nested_or_ancestor_destinations(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - source = root / "parent" / "source" - source.mkdir(parents=True) - skill_file = source / "SKILL.md" - skill_file.write_text("name: distilly\n", encoding="utf-8") - nested = source / ".agents" / "skills" / "distilly" - - for installer in (install_openclaw_skill, install_codex_skill): - with self.assertRaisesRegex(ValueError, "must not overlap"): - installer(source, nested, force=True) - with self.assertRaisesRegex(ValueError, "must not overlap"): - installer(source, source.parent, force=True) - - self.assertTrue(skill_file.exists()) - self.assertFalse(nested.exists()) - - def test_openclaw_generated_skill_installer_writes_host_skill_folder(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - generated_root = tmp_root / "skills" / "relationship" - openclaw_skills = tmp_root / ".openclaw" / "workspace" / "skills" - - skill_dir = skill_writer.create_skill( - generated_root, - "mireille", - { - "character": "relationship", - "name": "Mireille", - }, - "Work body", - "Persona body", - ) - - result = install_openclaw_generated_skill( - skill_dir, - openclaw_skills, - force=True, - ) - - installed_file = openclaw_skills / "relationship-mireille" / "SKILL.md" - metadata_file = openclaw_skills / "relationship-mireille" / ".distilly-install.json" - - self.assertEqual(result["command_name"], "relationship-mireille") - self.assertTrue(installed_file.exists()) - self.assertTrue(metadata_file.exists()) - self.assertIn( - "name: relationship-mireille", - installed_file.read_text(encoding="utf-8"), - ) - - def test_codex_generated_skill_installer_writes_host_skill_folder(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - tmp_root = Path(tmp_dir) - generated_root = tmp_root / "skills" / "celebrity" - codex_skills = tmp_root / ".agents" / "skills" - - skill_dir = skill_writer.create_skill( - generated_root, - "zhou-qimo", - { - "character": "celebrity", - "name": "周奇墨", - "classification": {"language": "zh-CN"}, - }, - "Work body", - "Persona body", - ) - - result = install_codex_generated_skill( - skill_dir, - codex_skills, - force=True, - ) - - installed_file = codex_skills / "celebrity-zhou-qimo" / "SKILL.md" - metadata_file = codex_skills / "celebrity-zhou-qimo" / ".distilly-install.json" - - self.assertEqual(result["command_name"], "celebrity-zhou-qimo") - self.assertTrue(installed_file.exists()) - self.assertTrue(metadata_file.exists()) - self.assertIn( - "name: celebrity-zhou-qimo", - installed_file.read_text(encoding="utf-8"), - ) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_research_tools.py b/tests/test_research_tools.py deleted file mode 100644 index 2f21dae6..00000000 --- a/tests/test_research_tools.py +++ /dev/null @@ -1,278 +0,0 @@ -from __future__ import annotations - -import tempfile -import unittest -from pathlib import Path -import sys - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from research.merge_research import merge_research # noqa: E402 -from research.quality_check import evaluate_skill_text # noqa: E402 -from research.srt_to_transcript import clean_subtitle_text, convert_file # noqa: E402 - - -class SubtitleTranscriptTest(unittest.TestCase): - def test_clean_subtitle_text_removes_timestamps_and_duplicates(self) -> None: - subtitle = """WEBVTT - -00:00:01.000 --> 00:00:03.000 -Hello - -00:00:03.000 --> 00:00:05.000 -Hello - -00:00:05.000 --> 00:00:07.000 align:start position:0% -World. -""" - transcript = clean_subtitle_text(subtitle) - self.assertNotIn("-->", transcript) - self.assertNotIn("", transcript) - self.assertIn("Hello World.", transcript) - - def test_convert_file_writes_transcript(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - input_path = Path(tmp_dir) / "sample.srt" - output_path = Path(tmp_dir) / "sample_transcript.txt" - input_path.write_text( - "1\n00:00:00,000 --> 00:00:01,000\nLine one.\n", - encoding="utf-8", - ) - convert_file(input_path, output_path) - self.assertIn("Line one.", output_path.read_text(encoding="utf-8")) - - -class ResearchMergeTest(unittest.TestCase): - def test_merge_research_writes_summary(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - skill_dir = Path(tmp_dir) / "celebrity" - raw_dir = skill_dir / "knowledge" / "research" / "raw" - raw_dir.mkdir(parents=True) - (raw_dir / "01.md").write_text( - "# Note\n" - "## Source Metadata\n" - "- URL: https://example.com/a\n" - "- Grounding level: primary\n" - "## Evidence\n" - "- Strong focus on first-person essays\n" - "## Contradictions\n" - "- Repeats a skeptical framing\n" - "## Inferences\n" - "- Likely values first-person observation.\n" - "primary source\n", - encoding="utf-8", - ) - (raw_dir / "02.md").write_text( - "# Note\n" - "## Source Metadata\n" - "- URL: https://example.com/b\n" - "- Grounding level: secondary\n" - "## Evidence\n" - "- Uses metaphor under pressure\n" - "## Contradictions\n" - "- Tension between compression and warmth\n" - "## Inferences\n" - "- Metaphor acts as a softening device.\n", - encoding="utf-8", - ) - - summary_path = merge_research(skill_dir) - summary = summary_path.read_text(encoding="utf-8") - - self.assertIn("# Research Summary", summary) - self.assertIn("Unique URLs: 2", summary) - self.assertIn("Total note chars:", summary) - self.assertIn("Potential long quote lines: 0", summary) - self.assertIn("Source metadata blocks: 2", summary) - self.assertIn("Contradiction bullets: 2", summary) - self.assertIn("Inference bullets: 2", summary) - self.assertIn("01.md", summary) - self.assertIn("Strong focus on first-person essays", summary) - - -class QualityCheckTest(unittest.TestCase): - def test_evaluate_skill_text_detects_expected_signals(self) -> None: - text = """ -## mental models -- First-principles reasoning -- Skeptical framing -- Long-horizon tradeoffs - -## intellectual genealogy -- Influenced by long-form systems thinking and adversarial debate. - -## agentic protocol -- Step 1: classify the question. -- Step 2: inspect known-answer anchors. -- Step 3: state uncertainty before extrapolation. - -## expression DNA -- Sentence rhythm is clipped. -- Uses metaphor when disagreeing. - -## honest boundaries -- States what they do not know. - -## contradictions -- Alternates between certainty and doubt. - -Sources: -https://example.com/articles/long-form-profile -https://example.com/interviews/episode-42 - -limitations: -- avoids operational detail -""" - report = evaluate_skill_text(text) - self.assertTrue(report["passed"]) - self.assertTrue(all(report["checks"].values())) - self.assertTrue(report["checks"]["copyright_safety"]) - - def test_source_grounding_rejects_generic_homepages(self) -> None: - text = """ -## mental models -- First-principles reasoning -- Skeptical framing -- Long-horizon tradeoffs - -## expression DNA -- Sentence rhythm is clipped. -- Uses metaphor when disagreeing. - -## honest boundaries -- States what they do not know. - -## contradictions -- Alternates between certainty and doubt. - -## limitations -- Needs more verified external material. - -Sources: -https://www.iqiyi.com/ -https://space.bilibili.com/ -https://www.zhihu.com/topic/ -""" - report = evaluate_skill_text(text) - self.assertFalse(report["checks"]["source_grounding"]) - self.assertEqual(report["grounded_url_count"], 0) - - def test_budget_unfriendly_profile_requires_deeper_research_metrics(self) -> None: - text = """ -## mental models -- First-principles reasoning -- Skeptical framing -- Long-horizon tradeoffs - -## expression DNA -- Sentence rhythm is clipped. -- Uses metaphor when disagreeing. - -## honest boundaries -- States what they do not know. - -## contradictions -- Alternates between certainty and doubt. - -## limitations -- Needs more verified external material. - -Sources: -https://example.com/articles/long-form-profile -https://example.com/interviews/episode-42 -https://example.com/interviews/episode-43 -https://example.com/talks/keynote-2019 -""" - report = evaluate_skill_text( - text, - profile="budget-unfriendly", - research_metrics={ - "files_scanned": 4, - "unique_urls": 4, - "primary_source_markers": 1, - "source_metadata_blocks": 4, - "contradiction_bullets": 2, - "inference_bullets": 2, - "long_quote_lines": 0, - "track_coverage_count": 4, - "research_audit_present": False, - "synthesis_review_present": False, - "validation_review_present": False, - "research_audit_pass": False, - "validation_review_pass": False, - "known_answer_questions": 1, - "edge_case_markers": 0, - }, - ) - self.assertFalse(report["passed"]) - self.assertFalse(report["checks"]["research_depth"]) - self.assertFalse(report["checks"]["review_chain"]) - self.assertFalse(report["checks"]["validation_depth"]) - - def test_budget_unfriendly_profile_passes_with_full_review_chain(self) -> None: - text = """ -## mental models -- First-principles reasoning -- Skeptical framing -- Long-horizon tradeoffs - -## intellectual genealogy -- Influenced by long-form systems thinking and adversarial debate. - -## agentic protocol -- Step 1: classify the question. -- Step 2: inspect known-answer anchors. -- Step 3: state uncertainty before extrapolation. - -## expression DNA -- Sentence rhythm is clipped. -- Uses metaphor when disagreeing. - -## honest boundaries -- States what they do not know. - -## contradictions -- Alternates between certainty and doubt. - -## limitations -- Needs more verified external material. - -Sources: -https://example.com/articles/long-form-profile -https://example.com/interviews/episode-42 -https://example.com/interviews/episode-43 -https://example.com/talks/keynote-2019 -""" - report = evaluate_skill_text( - text, - profile="budget-unfriendly", - research_metrics={ - "files_scanned": 6, - "unique_urls": 8, - "primary_source_markers": 3, - "source_metadata_blocks": 6, - "contradiction_bullets": 6, - "inference_bullets": 6, - "long_quote_lines": 0, - "track_coverage_count": 6, - "high_tier_sources": 4, - "mid_tier_sources": 2, - "low_tier_sources": 1, - "weighted_source_primary_ratio": 67, - "research_audit_present": True, - "synthesis_review_present": True, - "validation_review_present": True, - "research_audit_pass": True, - "validation_review_pass": True, - "known_answer_questions": 2, - "edge_case_markers": 1, - }, - ) - self.assertTrue(report["passed"]) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_skill_entrypoint_docs.py b/tests/test_skill_entrypoint_docs.py deleted file mode 100644 index 767eeb90..00000000 --- a/tests/test_skill_entrypoint_docs.py +++ /dev/null @@ -1,328 +0,0 @@ -from pathlib import Path -import re -import unittest - - -ROOT = Path(__file__).resolve().parents[1] -SUPPORTED_HOSTS = ( - "Claude Code", - "Hermes Agent", - "OpenClaw", - "Codex", - "DeepSeek Harness", - "Pi coding agent", - "Grok Build", - "OpenCode", -) -HOST_LOGO_FILES = ( - "claude-code-wordmark-dark.svg", - "claude-code-wordmark-light.svg", - "hermes-agent-wordmark.png", - "openclaw-wordmark-dark.svg", - "openclaw-wordmark-light.svg", - "codex-mark-dark.png", - "codex-mark-light.png", - "deepseek-wordmark-dark.svg", - "deepseek-wordmark-light.svg", - "pi-mark.svg", - "grok-build-mark-dark.png", - "grok-build-mark-light.png", - "opencode-wordmark-dark.svg", - "opencode-wordmark-light.svg", -) -# Slash invocations start at a text boundary; slashes inside install paths are allowed. -README_INVOCATION_PATTERNS = ( - re.compile(r"\$distilly\b"), - re.compile(r"@distilly\b"), - re.compile(r"(? None: - self.assertIn("### 3️⃣", content, f"missing host section in {source}") - host_section = content.split("### 3️⃣", 1)[1].split("\n---", 1)[0] - host_rows = re.findall(r']*alt="([^"]+)"', host_section) - - self.assertEqual(list(SUPPORTED_HOSTS), host_rows, f"wrong host list in {source}") - self.assertNotIn("🟣 **Claude Code**", host_section, f"text host rows remain in {source}") - self.assertNotIn( - "img.shields.io/badge/Claude%20Code-Skill", - content, - f"duplicate host badges remain in {source}", - ) - logo_prefix = "docs/assets/hosts/" if source == "README.md" else "../assets/hosts/" - for logo_file in HOST_LOGO_FILES: - self.assertIn(logo_prefix + logo_file, host_section, f"missing {logo_file} in {source}") - self.assertIn("Grok Bot", host_section, f"missing Grok Bot preview in {source}") - self.assertIn("{character}-{slug}", content, f"missing generated Skill name in {source}") - self.assertNotIn("## 📂", content, f"project structure section remains in {source}") - for pattern in README_INVOCATION_PATTERNS: - self.assertIsNone( - pattern.search(content), - f"invocation tutorial matching {pattern.pattern!r} remains in {source}", - ) - - def _assert_readme_quick_start( - self, content: str, source: str, install_link: str - ) -> None: - self.assertIn("## ⚡", content, f"missing install section in {source}") - quick_start = content.split("## ⚡", 1)[1].split("\n## ✨", 1)[0] - self.assertIn("### 🤖", quick_start, f"missing Agent path in {source}") - self.assertIn("### 👤", quick_start, f"missing human path in {source}") - self.assertIn( - "git clone https://github.com/titanwings/distilly " - "", - quick_start, - f"missing short clone command in {source}", - ) - self.assertIn(install_link, quick_start, f"missing install guide link in {source}") - self.assertIn("{character}-{slug}", quick_start) - self.assertNotIn("
", quick_start, f"collapsed install table remains in {source}") - self.assertNotIn( - "tools/install_generated_skill.py", - quick_start, - f"generated Skill installer detail remains in {source}", - ) - self.assertNotIn( - "tools/research/", - quick_start, - f"celebrity research commands remain in {source}", - ) - - def test_root_skill_uses_distilly_entrypoint(self) -> None: - content = (ROOT / "SKILL.md").read_text(encoding="utf-8") - self.assertIn("name: distilly", content) - self.assertIn("`/distilly`", content) - self.assertIn("`$distilly`", content) - self.assertIn("`/skill:distilly`", content) - self.assertNotIn("name: dot-skill", content) - self.assertNotIn("`/dot-skill`", content) - self.assertIn("兼容宿主", content) - self.assertIn("compatible hosts", content.lower()) - self.assertIn("Grok Build", content) - self.assertIn("Grok Bot", content) - self.assertIn("Pi coding agent", content) - self.assertIn("管理操作", content) - self.assertIn("tools/skill_writer.py", content) - self.assertIn("tools/research/xquik_public_posts.py", content) - self.assertIn("prompts/celebrity/research.md", content) - self.assertIn("budget-unfriendly", content) - self.assertIn("references/celebrity_budget_unfriendly_framework.md", content) - self.assertIn("01_core_profile.md", content) - self.assertIn("03_expression_and_reception.md", content) - self.assertIn("Files scanned >= 3", content) - self.assertIn("Unique URLs >= 2", content) - self.assertIn("Potential long quote lines = 0", content) - self.assertIn("实际打开过的具体页面", content) - self.assertIn("actual inspected pages", content) - self.assertIn("01_writings.md", content) - self.assertIn("06_timeline.md", content) - self.assertIn("Files scanned >= 6", content) - self.assertIn("Unique URLs >= 8", content) - self.assertIn("Primary-source markers >= 3", content) - self.assertIn("research_audit.md", content) - self.assertIn("--work-patch /tmp/distilly_{slug}_work_patch.md", content) - self.assertIn("Do not hand-edit `work.md`", content) - self.assertIn("{distilly_skill_root}", content) - self.assertIn("${CLAUDE_SKILL_DIR}", content) - self.assertIn("Do not assume the shell's current working directory", content) - self.assertNotIn("python3 tools/", content) - self.assertNotIn("`/list-skills`", content) - self.assertNotIn("Compatibility aliases:", content) - - def test_readme_quick_start_and_install_detail_contract(self) -> None: - readme = (ROOT / "README.md").read_text(encoding="utf-8") - install = (ROOT / "INSTALL.md").read_text(encoding="utf-8") - install_en = (ROOT / "INSTALL_EN.md").read_text(encoding="utf-8") - skill = (ROOT / "SKILL.md").read_text(encoding="utf-8") - - self._assert_readme_support_contract(readme, "README.md") - self._assert_readme_quick_start(readme, "README.md", "(INSTALL_EN.md)") - self.assertNotIn("/dot-skill", readme) - self.assertNotIn("skills/dot-skill", readme) - self.assertIn("https://github.com/titanwings/distilly", readme) - self.assertNotIn("https://github.com/titanwings/colleague-skill", readme) - - self.assertIn(".claude/skills/distilly", install) - self.assertIn("~/.openclaw/workspace/skills/distilly", install) - self.assertIn("~/.agents/skills/distilly", install) - self.assertIn("~/.dsh/skills/distilly", install) - self.assertIn("~/.pi/agent/skills/distilly", install) - self.assertIn("~/.grok/skills/distilly", install) - self.assertIn("~/.config/opencode/skills/distilly", install) - self.assertIn(".opencode/skills/distilly", install) - self.assertIn("/distilly", install) - self.assertNotIn("/dot-skill", install) - self.assertNotIn("skills/dot-skill", install) - self.assertIn("https://github.com/titanwings/distilly", install) - self.assertNotIn("https://github.com/titanwings/colleague-skill", install) - self.assertIn("./skills/colleague", install) - self.assertIn("install_claude_generated_skill.py", install) - self.assertIn("install_openclaw_generated_skill.py", install) - self.assertIn("install_codex_generated_skill.py", install) - self.assertIn("install_openclaw_skill.py", install) - self.assertIn("install_codex_skill.py", install) - self.assertIn("tools/research/quality_check.py", install) - self.assertIn("tools/research/xquik_public_posts.py", install) - self.assertIn("tools/research/download_subtitles.sh", install) - self.assertIn("tools/research/merge_research.py", install) - self.assertIn("pip3 install -r requirements.txt", install) - self.assertIn("pip3 install yt-dlp", install) - self.assertIn("~/.claude/skills/distilly", install_en) - self.assertIn("~/.openclaw/workspace/skills/distilly", install_en) - self.assertIn("~/.hermes/skills/openclaw-imports/distilly", install_en) - self.assertIn("~/.agents/skills/distilly", install_en) - self.assertIn("~/.dsh/skills/distilly", install_en) - self.assertIn("~/.pi/agent/skills/distilly", install_en) - self.assertIn("~/.grok/skills/distilly", install_en) - self.assertIn("~/.config/opencode/skills/distilly", install_en) - self.assertIn("tools/install_generated_skill.py", install_en) - self.assertIn("tools/research/download_subtitles.sh", install_en) - self.assertIn("tools/research/srt_to_transcript.py", install_en) - self.assertIn("tools/research/merge_research.py", install_en) - self.assertIn("tools/research/quality_check.py", install_en) - self.assertIn("pip3 install -r requirements.txt", install_en) - self.assertIn("pip3 install yt-dlp", install_en) - self.assertIn("/{character}-{slug}", install) - self.assertIn("./skills/colleague", skill) - self.assertIn("DeepSeek Harness", skill) - self.assertIn("OpenCode", skill) - self.assertIn("eight agent hosts", readme.lower()) - self.assertIn("兼容宿主", install) - - for logo_file in HOST_LOGO_FILES: - self.assertTrue((ROOT / "docs" / "assets" / "hosts" / logo_file).is_file()) - - def test_repo_examples_live_under_skills_colleague(self) -> None: - self.assertTrue((ROOT / "skills" / "colleague" / "example_zhangsan").exists()) - self.assertTrue((ROOT / "skills" / "colleague" / "example_tianyi").exists()) - self.assertTrue((ROOT / "skills" / "colleague" / "example_jiaxiu").exists()) - self.assertFalse((ROOT / "colleagues").exists()) - - def test_multilingual_readmes_list_hosts_without_invocation_tutorials(self) -> None: - for readme_path in (ROOT / "docs" / "lang").glob("README_*.md"): - content = readme_path.read_text(encoding="utf-8") - self.assertNotIn("/dot-skill", content, f"stale /dot-skill in {readme_path.name}") - self._assert_readme_support_contract(content, readme_path.name) - self._assert_readme_quick_start( - content, - readme_path.name, - "(../../INSTALL.md)" - if readme_path.name == "README_ZH.md" - else "(../../INSTALL_EN.md)", - ) - self.assertIn( - "https://github.com/titanwings/distilly", - content, - f"missing published repository URL in {readme_path.name}", - ) - self.assertNotIn( - "https://github.com/titanwings/colleague-skill", - content, - f"stale repository URL in {readme_path.name}", - ) - self.assertNotIn( - "colleague_skill.pdf", - content, - f"broken local paper link in {readme_path.name}", - ) - - def test_non_chinese_docs_call_the_product_lark(self) -> None: - non_chinese = [ROOT / "README.md", ROOT / "docs" / "lang" / "README_EN.md"] - non_chinese.extend( - ROOT / "docs" / "lang" / f"README_{language}.md" - for language in ("DE", "ES", "JA", "KO", "PT", "RU") - ) - for path in non_chinese: - content = path.read_text(encoding="utf-8") - self.assertIn("Lark", content, f"missing Lark in {path.name}") - self.assertNotIn("Feishu", content, f"stale visible Feishu name in {path.name}") - - chinese = (ROOT / "docs" / "lang" / "README_ZH.md").read_text(encoding="utf-8") - self.assertIn("飞书", chinese) - - def test_non_chinese_surfaces_do_not_mix_chinese_copy(self) -> None: - single_language_paths = [ - ROOT / "README.md", - ROOT / "ROADMAP.md", - ROOT / "CONTRIBUTING.md", - ROOT / "CITATION.cff", - ROOT / "INSTALL_EN.md", - ROOT / "docs" / "lang" / "README_EN.md", - ROOT / "prompts" / "celebrity" / "research.md", - ROOT / "prompts" / "celebrity" / "budget_unfriendly" / "research.md", - ROOT / "references" / "celebrity_budget_unfriendly_framework.md", - ROOT / ".github" / "PULL_REQUEST_TEMPLATE.md", - ] - single_language_paths.extend( - ROOT / "docs" / "lang" / f"{stem}_{language}.md" - for stem in ("README", "ROADMAP") - for language in ("DE", "ES", "KO", "PT", "RU") - ) - single_language_paths.extend( - ROOT / ".github" / "ISSUE_TEMPLATE" / name - for name in ( - "bug_report.md", - "config.yml", - "feature_request.md", - "question.md", - ) - ) - - for path in single_language_paths: - content = path.read_text(encoding="utf-8") - self.assertIsNone( - HAN_CHARACTER.search(content), - f"Chinese copy remains in {path.relative_to(ROOT)}", - ) - - skill = (ROOT / "SKILL.md").read_text(encoding="utf-8") - frontmatter = skill.split("---", 2)[1] - chinese_skill = skill.split("# English Version", 1)[0] - english_skill = skill.split("# English Version", 1)[1] - self.assertIsNone(HAN_CHARACTER.search(frontmatter)) - self.assertIsNone(HAN_CHARACTER.search(english_skill)) - self.assertIn("本 Skill 支持中英文", chinese_skill) - self.assertIsNotNone(HAN_CHARACTER.search(chinese_skill)) - - for name in ("README_JA.md", "ROADMAP_JA.md"): - content = (ROOT / "docs" / "lang" / name).read_text(encoding="utf-8") - for stale_phrase in ("原同事", "留痕", "企業微信"): - self.assertNotIn(stale_phrase, content, f"Chinese phrase in {name}") - - def test_code_uses_new_names_with_explicit_legacy_fallbacks(self) -> None: - writer = (ROOT / "tools" / "skill_writer.py").read_text(encoding="utf-8") - schema = (ROOT / "tools" / "skill_schema.py").read_text(encoding="utf-8") - codex_installer = (ROOT / "tools" / "install_codex_skill.py").read_text(encoding="utf-8") - generated_installer = ( - ROOT / "tools" / "install_generated_skill_common.py" - ).read_text(encoding="utf-8") - - self.assertIn("DISTILLY_AUTO_INSTALL_CLAUDE", writer) - self.assertIn("DOT_SKILL_AUTO_INSTALL_CLAUDE", writer) - self.assertIn('engine.setdefault("name", "distilly")', schema) - self.assertIn('generation.setdefault("engine", "distilly")', schema) - self.assertIn('Path.home() / ".agents" / "skills" / "distilly"', codex_installer) - self.assertIn('.distilly-install.json', generated_installer) - - for collector_name in ( - "feishu_auto_collector.py", - "feishu_mcp_client.py", - "dingtalk_auto_collector.py", - "slack_auto_collector.py", - ): - collector = (ROOT / "tools" / collector_name).read_text(encoding="utf-8") - self.assertIn('Path.home() / ".distilly"', collector) - self.assertIn('Path.home() / ".colleague-skill"', collector) - self.assertIn("CONFIG_PATH.chmod(0o600)", collector) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_skill_writer.py b/tests/test_skill_writer.py deleted file mode 100644 index df738340..00000000 --- a/tests/test_skill_writer.py +++ /dev/null @@ -1,719 +0,0 @@ -from __future__ import annotations - -import json -import os -import sys -import tempfile -import unittest -from pathlib import Path - - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -import skill_writer # noqa: E402 -import version_manager # noqa: E402 -from skill_presets import ( # noqa: E402 - get_character_preset, - get_research_profile_preset, - resolve_existing_storage_root, -) -from skill_schema import validate_path_segment # noqa: E402 - - -class SkillWriterTest(unittest.TestCase): - def test_slugify_produces_portable_kebab_case(self) -> None: - self.assertEqual(skill_writer.slugify("Zadie Smith"), "zadie-smith") - self.assertEqual(skill_writer.slugify("Élodie"), "elodie") - self.assertEqual(skill_writer.slugify("A/B"), "a-b") - - def test_create_skill_rejects_unsafe_slug_before_writing(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - with self.assertRaisesRegex(ValueError, "kebab-case"): - skill_writer.create_skill( - root / "skills" / "colleague", - "../escape", - {"name": "Unsafe"}, - "Work body", - "Persona body", - ) - self.assertFalse((root / "skills" / "escape").exists()) - - def test_legacy_path_segments_are_windows_safe(self) -> None: - self.assertEqual(validate_path_segment("Zadie Smith"), "Zadie Smith") - self.assertEqual(validate_path_segment("Élodie"), "Élodie") - for value in ("C:", "foo:bar", "CON", "nul.txt", "trailing.", "trailing "): - with self.subTest(value=value): - with self.assertRaisesRegex(ValueError, "safe path segment"): - validate_path_segment(value) - - def test_create_colleague_uses_portable_names_and_adds_engine_schema(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "colleague" - meta = { - "name": "Eulalie", - "profile": { - "company": "ByteDance", - "level": "L2-1", - "role": "Backend Engineer", - "mbti": "INTJ", - }, - "tags": { - "personality": ["direct", "data-driven"], - "culture": ["byte-dance-style"], - }, - } - - skill_dir = skill_writer.create_skill( - base_dir, - "zhangsan", - meta, - "Work body", - "Persona body", - ) - - saved_meta = json.loads( - (skill_dir / "meta.json").read_text(encoding="utf-8") - ) - manifest = json.loads((skill_dir / "manifest.json").read_text(encoding="utf-8")) - combined_skill = (skill_dir / "SKILL.md").read_text(encoding="utf-8") - work_skill = (skill_dir / "work_skill.md").read_text(encoding="utf-8") - persona_skill = (skill_dir / "persona_skill.md").read_text(encoding="utf-8") - - self.assertEqual(saved_meta["schema_version"], "3") - self.assertEqual(saved_meta["kind"], "meta-skill") - self.assertEqual(saved_meta["character"], "colleague") - self.assertEqual(saved_meta["preset"], "distilly.colleague.v1") - self.assertEqual(saved_meta["engine"]["name"], "distilly") - self.assertEqual(saved_meta["generation"]["engine"], "distilly") - self.assertEqual(saved_meta["type"], "colleague") - self.assertEqual(saved_meta["id"], "meta-skill.colleague.zhangsan") - self.assertEqual(saved_meta["artifacts"]["combined_name"], "colleague-zhangsan") - self.assertEqual(saved_meta["artifacts"]["combined_command"], "colleague-zhangsan") - self.assertEqual(saved_meta["compat"]["legacy_command"], "/create-colleague") - self.assertEqual(manifest["kind"], "meta-skill") - self.assertEqual(manifest["character"], "colleague") - self.assertEqual(manifest["preset"], "distilly.colleague.v1") - self.assertEqual(manifest["install"]["slash_commands"]["default"], "colleague-zhangsan") - self.assertEqual( - manifest["install"]["compatible_runtimes"], - [ - "claude-code", - "openclaw", - "hermes", - "codex", - "deepseek-harness", - "grok-build", - "pi", - "opencode", - ], - ) - self.assertEqual( - manifest["install"]["installers"]["openclaw"], - "tools/install_openclaw_generated_skill.py", - ) - self.assertEqual( - manifest["install"]["installers"]["codex"], - "tools/install_codex_generated_skill.py", - ) - self.assertIn("name: colleague-zhangsan", combined_skill) - self.assertIn("## PART A: Work", combined_skill) - self.assertIn("name: colleague-zhangsan-work", work_skill) - self.assertIn("work capability only", work_skill) - self.assertIn("name: colleague-zhangsan-persona", persona_skill) - self.assertIn("persona only", persona_skill) - - def test_create_relationship_uses_character_preset_metadata(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "relationship" - meta = { - "character": "relationship", - "name": "Mireille", - "profile": { - "role": "Designer", - }, - } - - skill_dir = skill_writer.create_skill( - base_dir, - "mireille", - meta, - "Work body", - "Persona body", - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - manifest = json.loads((skill_dir / "manifest.json").read_text(encoding="utf-8")) - combined_skill = (skill_dir / "SKILL.md").read_text(encoding="utf-8") - - self.assertEqual(saved_meta["kind"], "meta-skill") - self.assertEqual(saved_meta["character"], "relationship") - self.assertEqual(saved_meta["preset"], "distilly.relationship.v1") - self.assertEqual(saved_meta["type"], "relationship") - self.assertEqual(saved_meta["classification"]["gallery_category"], "Relationship") - self.assertEqual(saved_meta["compat"]["legacy_storage_root"], "skills/relationship") - self.assertEqual(manifest["id"], "meta-skill.relationship.mireille") - self.assertEqual(manifest["character"], "relationship") - self.assertEqual(saved_meta["artifacts"]["combined_command"], "relationship-mireille") - self.assertIn("name: relationship-mireille", combined_skill) - - def test_create_skill_renders_chinese_chrome_when_language_is_zh_cn(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "relationship" - meta = { - "character": "relationship", - "name": "Mireille", - "classification": { - "language": "zh-CN", - }, - } - - skill_dir = skill_writer.create_skill( - base_dir, - "mireille", - meta, - "Work body", - "Persona body", - ) - - combined_skill = (skill_dir / "SKILL.md").read_text(encoding="utf-8") - work_skill = (skill_dir / "work_skill.md").read_text(encoding="utf-8") - persona_skill = (skill_dir / "persona_skill.md").read_text(encoding="utf-8") - - self.assertIn("## PART A:工作能力", combined_skill) - self.assertIn("运行规则", combined_skill) - self.assertIn("仅 Work,无 Persona", work_skill) - self.assertIn("仅 Persona,无工作能力", persona_skill) - - def test_work_only_skill_replaces_persona_handoff(self) -> None: - zh_handoff = "如果被问到职责范围外的问题,以该同事的方式回应(参见 Persona 部分)。" - en_handoff = ( - "If you are asked a question outside your recorded responsibilities, " - "respond in this colleague's style (see the Persona section)." - ) - zh_work_content = ( - "## 工作能力使用说明\n\n" - "当用户要求你完成以下任务时,严格按照上述规范执行。\n\n" - f"{zh_handoff}\n" - ) - en_work_content = ( - "## Scope rule\n\n" - "If asked outside your recorded responsibilities:\n" - "- State the evidence gap\n\n" - "## Persona naming note\n\n" - "Keep this documentation sentence.\n\n" - f"{en_handoff}\n" - ) - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "colleague" - zh_meta = { - "name": "Eulalie", - "language": "zh-CN", - "profile": { - "company": "ByteDance", - "level": "L2-1", - "role": "Backend Engineer", - }, - } - en_meta = { - "name": "Eulalie", - "language": "en", - "profile": { - "company": "ByteDance", - "level": "L2-1", - "role": "Backend Engineer", - }, - } - - zh_dir = skill_writer.create_skill( - base_dir / "zh", - "zhangsan", - zh_meta, - zh_work_content, - "Persona body", - ) - en_dir = skill_writer.create_skill( - base_dir / "en", - "zhangsan", - en_meta, - en_work_content, - "Persona body", - ) - - zh_stored_work = (zh_dir / "work.md").read_text(encoding="utf-8") - zh_combined = (zh_dir / "SKILL.md").read_text(encoding="utf-8") - en_stored_work = (en_dir / "work.md").read_text(encoding="utf-8") - en_combined = (en_dir / "SKILL.md").read_text(encoding="utf-8") - zh_work_skill = (zh_dir / "work_skill.md").read_text(encoding="utf-8") - en_work_skill = (en_dir / "work_skill.md").read_text(encoding="utf-8") - - self.assertIn(zh_handoff, zh_stored_work) - self.assertIn(zh_handoff, zh_combined) - self.assertIn(en_handoff, en_stored_work) - self.assertIn(en_handoff, en_combined) - self.assertNotIn(zh_handoff, zh_work_skill) - self.assertNotIn(en_handoff, en_work_skill) - self.assertIn("If asked outside your recorded responsibilities:", en_work_skill) - self.assertIn("## Persona naming note", en_work_skill) - self.assertIn("Keep this documentation sentence.", en_work_skill) - self.assertIn(skill_writer.WORK_ONLY_FALLBACK_ZH, zh_work_skill) - self.assertIn(skill_writer.WORK_ONLY_FALLBACK_EN, en_work_skill) - self.assertIn("不要臆造缺失信息", zh_work_skill) - self.assertNotIn("不要推断", zh_work_skill) - self.assertIn("Do not fabricate missing information", en_work_skill) - self.assertNotIn("Do not infer", en_work_skill) - self.assertNotIn(skill_writer.WORK_ONLY_FALLBACK_ZH, zh_combined) - self.assertNotIn(skill_writer.WORK_ONLY_FALLBACK_EN, en_combined) - - def test_create_celebrity_adds_research_dirs_and_toolchain(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "celebrity" - meta = { - "character": "celebrity", - "name": "Zadie Smith", - "profile": { - "identity": "Novelist", - "known_for": "Essay and criticism", - }, - "tags": ["literature", "essay", "public-intellectual"], - "knowledge_sources": ["interview", "essay"], - } - - skill_dir = skill_writer.create_skill( - base_dir, - "zadie-smith", - meta, - "Work body", - "Persona body", - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - manifest = json.loads((skill_dir / "manifest.json").read_text(encoding="utf-8")) - - self.assertEqual(saved_meta["character"], "celebrity") - self.assertEqual(saved_meta["preset"], "distilly.celebrity.v1") - self.assertEqual(saved_meta["research_profile"], "budget-friendly") - self.assertIn("research_tools", saved_meta["engine"]) - self.assertEqual(saved_meta["engine"]["research_profile"], "budget-friendly") - self.assertIn("research_tools", manifest["toolchain"]) - self.assertEqual(manifest["research_profile"], "budget-friendly") - self.assertEqual( - saved_meta["classification"]["tags"], - ["literature", "essay", "public-intellectual"], - ) - self.assertIn("Novelist", saved_meta["summary"]) - self.assertIn("Essay and criticism", saved_meta["summary"]) - self.assertTrue((skill_dir / "knowledge" / "research" / "raw").exists()) - self.assertTrue((skill_dir / "knowledge" / "research" / "merged").exists()) - self.assertTrue((skill_dir / "knowledge" / "transcripts").exists()) - self.assertTrue((skill_dir / "knowledge" / "subtitles").exists()) - - def test_create_celebrity_budget_unfriendly_embeds_profile_config(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "celebrity" - meta = { - "character": "celebrity", - "research_profile": "budget-unfriendly", - "name": "Xu Zhisheng", - "classification": {"language": "zh-CN"}, - } - - skill_dir = skill_writer.create_skill( - base_dir, - "xu-zhisheng", - meta, - "Work body", - "Persona body", - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - manifest = json.loads((skill_dir / "manifest.json").read_text(encoding="utf-8")) - - self.assertEqual(saved_meta["research_profile"], "budget-unfriendly") - self.assertEqual(saved_meta["engine"]["quality_profile"], "budget-unfriendly") - self.assertIn( - "prompts/celebrity/budget_unfriendly/research.md", - saved_meta["engine"]["research_profile_bundle"].values(), - ) - self.assertIn( - "prompts/celebrity/budget_unfriendly/audit.md", - saved_meta["engine"]["research_profile_bundle"].values(), - ) - self.assertIn( - "references/celebrity_budget_unfriendly_framework.md", - saved_meta["engine"]["research_profile_references"], - ) - self.assertEqual(manifest["research_profile"], "budget-unfriendly") - self.assertEqual(manifest["toolchain"]["quality_profile"], "budget-unfriendly") - self.assertEqual(manifest["toolchain"]["merge_strategy"], "deep") - - def test_create_celebrity_accepts_string_profile_from_runtime_meta(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "celebrity" - meta = { - "character": "celebrity", - "name": "徐志胜", - "display_name": "徐志胜", - "classification": {"language": "zh-CN"}, - "profile": "中国脱口秀演员,以自嘲式观察喜剧著称。", - } - - skill_dir = skill_writer.create_skill( - base_dir, - "xu-zhisheng", - meta, - "Work body", - "Persona body", - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - self.assertEqual(saved_meta["profile"], "中国脱口秀演员,以自嘲式观察喜剧著称。") - self.assertIn("中国脱口秀演员", saved_meta["summary"]) - - def test_existing_dot_skill_metadata_keeps_legacy_engine_identifiers(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "colleague" - skill_dir = skill_writer.create_skill( - base_dir, - "legacy", - { - "name": "Legacy", - "preset": "dot.colleague.v1", - "engine": {"name": "dot-skill"}, - "generation": {"engine": "dot-skill"}, - "artifacts": { - "combined_name": "colleague_legacy", - "work_name": "colleague_legacy_work", - "persona_name": "colleague_legacy_persona", - }, - }, - "Work body", - "Persona body", - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - self.assertEqual(saved_meta["preset"], "dot.colleague.v1") - self.assertEqual(saved_meta["engine"]["name"], "dot-skill") - self.assertEqual(saved_meta["generation"]["engine"], "dot-skill") - self.assertEqual(saved_meta["artifacts"]["combined_name"], "colleague_legacy") - self.assertIn( - "name: colleague_legacy", - (skill_dir / "SKILL.md").read_text(encoding="utf-8"), - ) - - def test_update_preserves_names_from_legacy_meta_without_artifacts(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - skill_dir = Path(tmp_dir) / "skills" / "colleague" / "legacy_person" - skill_dir.mkdir(parents=True) - (skill_dir / "versions").mkdir() - (skill_dir / "meta.json").write_text( - json.dumps( - { - "name": "Legacy Person", - "type": "colleague", - "version": "v1", - } - ), - encoding="utf-8", - ) - (skill_dir / "work.md").write_text("Legacy work\n", encoding="utf-8") - (skill_dir / "persona.md").write_text("Legacy persona\n", encoding="utf-8") - legacy_names = { - "SKILL.md": "colleague_legacy_person", - "work_skill.md": "colleague_legacy_person_work", - "persona_skill.md": "colleague_legacy_person_persona", - } - for filename, name in legacy_names.items(): - (skill_dir / filename).write_text( - f"---\nname: {name}\ndescription: Legacy\n---\n\nLegacy body\n", - encoding="utf-8", - ) - - skill_writer.update_skill(skill_dir, work_patch="Updated work") - - for filename, name in legacy_names.items(): - content = (skill_dir / filename).read_text(encoding="utf-8") - self.assertIn(f"name: {name}", content) - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - self.assertEqual( - saved_meta["artifacts"]["combined_command"], - "colleague-legacy-person", - ) - - def test_update_rejects_traversal_in_stored_version_before_backup(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - root = Path(tmp_dir) - skill_dir = skill_writer.create_skill( - root / "skills" / "colleague", - "unsafe-version", - {"name": "Unsafe Version"}, - "Work body", - "Persona body", - ) - meta_path = skill_dir / "meta.json" - meta = json.loads(meta_path.read_text(encoding="utf-8")) - meta["version"] = "../../../../escape" - meta["lifecycle"]["version"] = "../../../../escape" - meta_path.write_text(json.dumps(meta), encoding="utf-8") - - with self.assertRaisesRegex(ValueError, "safe path segment"): - skill_writer.update_skill(skill_dir, work_patch="Should not be written") - - self.assertFalse((root / "escape").exists()) - self.assertNotIn( - "Should not be written", - (skill_dir / "work.md").read_text(encoding="utf-8"), - ) - - def test_update_regenerates_manifest_and_archives_artifacts(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "colleague" - skill_dir = skill_writer.create_skill( - base_dir, - "zhangsan", - {"name": "Eulalie"}, - "Initial work", - "Initial persona", - ) - - new_version = skill_writer.update_skill( - skill_dir, - work_patch="More work", - correction={"scene": "challenged", "wrong": "apologize", "correct": "ask for evidence"}, - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - manifest = json.loads((skill_dir / "manifest.json").read_text(encoding="utf-8")) - archived_manifest = skill_dir / "versions" / "v1" / "manifest.json" - persona_doc = (skill_dir / "persona.md").read_text(encoding="utf-8") - - self.assertEqual(new_version, "v2") - self.assertEqual(saved_meta["version"], "v2") - self.assertEqual(saved_meta["corrections_count"], 1) - self.assertTrue(archived_manifest.exists()) - self.assertEqual(manifest["entrypoints"]["default"], "SKILL.md") - self.assertIn("apologize", persona_doc) - self.assertIn("ask for evidence", persona_doc) - - def test_update_accepts_multiple_persona_corrections_in_one_payload(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "celebrity" - skill_dir = skill_writer.create_skill( - base_dir, - "zhou-qimo", - { - "character": "celebrity", - "name": "周奇墨", - "classification": {"language": "zh-CN"}, - }, - "Initial work", - "Initial persona", - ) - - new_version = skill_writer.update_skill( - skill_dir, - correction={ - "persona_corrections": [ - { - "scene": "铺陈处境时", - "wrong": "一上来就下判断", - "correct": "先把处境讲得很普通,再轻轻点一下", - }, - { - "scene": "表达立场时", - "wrong": "写成明显自嘲型", - "correct": "和观众一起承认大家都在局里", - }, - ] - }, - ) - - saved_meta = json.loads((skill_dir / "meta.json").read_text(encoding="utf-8")) - persona_doc = (skill_dir / "persona.md").read_text(encoding="utf-8") - - self.assertEqual(new_version, "v2") - self.assertEqual(saved_meta["corrections_count"], 2) - self.assertIn("一上来就下判断", persona_doc) - self.assertIn("写成明显自嘲型", persona_doc) - self.assertEqual(persona_doc.count("## Correction Log"), 1) - - def test_update_replaces_existing_markdown_sections_instead_of_appending_duplicates(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "celebrity" - skill_dir = skill_writer.create_skill( - base_dir, - "zhou-qimo", - { - "character": "celebrity", - "name": "周奇墨", - "classification": {"language": "zh-CN"}, - }, - "\n".join( - [ - "# Work", - "", - "## 表达规范", - "", - "- 原始表述", - "", - "## 输出风格", - "", - "- 原始结构", - ] - ), - "\n".join( - [ - "# Persona", - "", - "## Layer 2: Expression DNA", - "", - "旧内容", - "", - "## Layer 3: Mental Models", - "", - "保持不变", - ] - ), - ) - - skill_writer.update_skill( - skill_dir, - work_patch="\n".join( - [ - "## 表达规范", - "", - "- 新的节奏控制", - "", - "## 输出风格", - "", - "- 新的结构模板", - ] - ), - persona_patch="\n".join( - [ - "## Layer 2: Expression DNA", - "", - "新内容", - ] - ), - ) - - work_doc = (skill_dir / "work.md").read_text(encoding="utf-8") - persona_doc = (skill_dir / "persona.md").read_text(encoding="utf-8") - - self.assertEqual(work_doc.count("## 表达规范"), 1) - self.assertEqual(work_doc.count("## 输出风格"), 1) - self.assertIn("新的节奏控制", work_doc) - self.assertNotIn("原始表述", work_doc) - self.assertEqual(persona_doc.count("## Layer 2: Expression DNA"), 1) - self.assertIn("新内容", persona_doc) - self.assertNotIn("旧内容", persona_doc) - - -class VersionManagerTest(unittest.TestCase): - def test_backup_and_rollback_include_manifest(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - base_dir = Path(tmp_dir) / "skills" / "colleague" - skill_dir = skill_writer.create_skill( - base_dir, - "zhangsan", - {"name": "Eulalie"}, - "v1 work", - "v1 persona", - ) - - version_manager.backup_current_version(skill_dir) - skill_writer.update_skill(skill_dir, work_patch="v2 work") - - success = version_manager.rollback(skill_dir, "v1") - restored_work = (skill_dir / "work.md").read_text(encoding="utf-8") - - self.assertTrue(success) - self.assertIn("v1 work", restored_work) - self.assertTrue((skill_dir / "versions" / "v1" / "manifest.json").exists()) - self.assertFalse(version_manager.rollback(skill_dir, "../v1")) - - def test_version_manager_can_still_resolve_legacy_colleagues_root(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - cwd = Path.cwd() - try: - os.chdir(tmp_dir) - legacy_base_dir = Path("colleagues") - skill_writer.create_skill( - legacy_base_dir, - "zhangsan", - {"name": "Eulalie"}, - "v1 work", - "v1 persona", - ) - - resolved = resolve_existing_storage_root("colleague", slug="zhangsan") - self.assertEqual(resolved, Path("colleagues")) - finally: - os.chdir(cwd) - - -class PromptPresetTest(unittest.TestCase): - def test_character_prompt_bundles_exist(self) -> None: - project_root = Path(__file__).resolve().parents[1] - - for character in ("colleague", "relationship", "celebrity"): - preset = get_character_preset(character) - for prompt_path in preset["prompt_bundle"].values(): - if not isinstance(prompt_path, str) or not prompt_path.startswith("prompts/"): - continue - self.assertTrue( - (project_root / prompt_path).exists(), - f"missing prompt file for {character}: {prompt_path}", - ) - for tool_path in preset.get("research_tools", {}).values(): - self.assertTrue( - (project_root / tool_path).exists(), - f"missing research tool for {character}: {tool_path}", - ) - for profile_name in preset.get("research_profiles", {}): - profile = get_research_profile_preset(character, profile_name) - for prompt_path in profile.get("prompt_bundle", {}).values(): - if not isinstance(prompt_path, str) or not prompt_path.startswith("prompts/"): - continue - self.assertTrue( - (project_root / prompt_path).exists(), - f"missing profile prompt file for {character}/{profile_name}: {prompt_path}", - ) - for reference_path in profile.get("references", []): - self.assertTrue( - (project_root / reference_path).exists(), - f"missing profile reference for {character}/{profile_name}: {reference_path}", - ) - - friendly_prompt = (project_root / "prompts" / "celebrity" / "research.md").read_text(encoding="utf-8") - self.assertIn("01_core_profile.md", friendly_prompt) - self.assertIn("03_expression_and_reception.md", friendly_prompt) - self.assertIn( - "do not collapse the whole pass into one monolithic note", - friendly_prompt.lower(), - ) - self.assertIn("actual inspected pages", friendly_prompt) - self.assertIn("tools/research/xquik_public_posts.py", friendly_prompt) - self.assertIn("untrusted candidate evidence", " ".join(friendly_prompt.split())) - - strict_prompt = ( - project_root - / "prompts" - / "celebrity" - / "budget_unfriendly" - / "research.md" - ).read_text(encoding="utf-8") - self.assertIn("01_writings.md", strict_prompt) - self.assertIn("06_timeline.md", strict_prompt) - self.assertIn("at least 8 grounded source URLs", strict_prompt) - self.assertIn("Do not replace these six files with one merged scratchpad", strict_prompt) - self.assertIn("actual inspected pages", strict_prompt) - self.assertIn("tools/research/xquik_public_posts.py", strict_prompt) - self.assertIn("untrusted candidate evidence", " ".join(strict_prompt.split())) - - -if __name__ == "__main__": - unittest.main() diff --git a/tests/test_xquik_public_posts.py b/tests/test_xquik_public_posts.py deleted file mode 100644 index da0d8d35..00000000 --- a/tests/test_xquik_public_posts.py +++ /dev/null @@ -1,242 +0,0 @@ -from __future__ import annotations - -import json -import os -import sys -import tempfile -import unittest -from pathlib import Path - -TOOLS_DIR = Path(__file__).resolve().parents[1] / "tools" -if str(TOOLS_DIR) not in sys.path: - sys.path.insert(0, str(TOOLS_DIR)) - -from research.xquik_public_posts import ( # noqa: E402 - API_CONTRACT, - API_URL, - MAX_CONTENT_CHARS, - CollectorError, - build_query, - collect_public_posts, - normalize_post, - validate_limit, - write_collection, -) - - -class FakeResponse: - def __init__(self, payload: object, status_code: int = 200) -> None: - self.payload = payload - self.status_code = status_code - - def json(self) -> object: - return self.payload - - -def sample_post(tweet_id: str = "123") -> dict: - return { - "id": tweet_id, - "text": "A public first-person observation.", - "createdAt": "2026-08-20T10:00:00Z", - "lang": "en", - "likeCount": 10, - "retweetCount": 2, - "replyCount": 1, - "quoteCount": 0, - "viewCount": 500, - "bookmarkCount": 3, - "author": { - "id": "42", - "username": "example_user", - "name": "Example User", - "verified": True, - }, - } - - -class QueryValidationTest(unittest.TestCase): - def test_build_query_accepts_username_with_at_prefix(self) -> None: - self.assertEqual(build_query(None, "@example_user"), "from:example_user") - - def test_build_query_rejects_invalid_username(self) -> None: - with self.assertRaisesRegex(CollectorError, "Invalid X username"): - build_query(None, "not-valid!") - - def test_validate_limit_rejects_unbounded_collection(self) -> None: - with self.assertRaisesRegex(CollectorError, "between 1 and 100"): - validate_limit(101) - - -class PostNormalizationTest(unittest.TestCase): - def test_normalize_post_canonicalizes_url_and_metrics(self) -> None: - post = normalize_post(sample_post()) - - self.assertIsNotNone(post) - self.assertEqual(post["url"], "https://x.com/example_user/status/123") - self.assertEqual(post["metrics"]["likes"], 10) - self.assertEqual(post["trust"], "untrusted_candidate_evidence") - - def test_normalize_post_truncates_long_form_text(self) -> None: - raw = sample_post() - raw["text"] = "x" * (MAX_CONTENT_CHARS + 50) - - post = normalize_post(raw) - - self.assertTrue(post["content_truncated"]) - self.assertEqual(len(post["content"]), MAX_CONTENT_CHARS) - - def test_normalize_post_accepts_current_normalized_contract(self) -> None: - raw = sample_post() - raw.pop("createdAt") - raw.pop("likeCount") - raw["created"] = 1_700_000_000 - raw["like_count"] = 11 - - post = normalize_post(raw) - - self.assertEqual(post["published_at"], "2023-11-14T22:13:20Z") - self.assertEqual(post["metrics"]["likes"], 11) - - def test_normalize_post_removes_terminal_control_characters(self) -> None: - raw = sample_post() - raw["text"] = "safe\x1b[31m text" - - post = normalize_post(raw) - - self.assertNotIn("\x1b", post["content"]) - - def test_normalize_post_rejects_non_x_source_url(self) -> None: - raw = sample_post() - raw["author"] = {} - raw["url"] = "https://example.com/example_user/status/123" - - self.assertIsNone(normalize_post(raw)) - - -class CollectionTest(unittest.TestCase): - def test_collect_public_posts_uses_one_bounded_read_request(self) -> None: - calls = [] - - def fake_get(url: str, **kwargs: object) -> FakeResponse: - calls.append((url, kwargs)) - return FakeResponse( - { - "tweets": [sample_post(), sample_post(), sample_post("456")], - "has_more": True, - "next_cursor": "opaque", - } - ) - - collection = collect_public_posts( - api_key="test-key", - query="from:example_user", - limit=20, - query_type="Latest", - subject="Example User", - request_get=fake_get, - collected_at="2026-08-21T00:00:00Z", - ) - - self.assertEqual(len(calls), 1) - self.assertEqual(calls[0][0], API_URL) - self.assertEqual(calls[0][1]["headers"]["x-api-key"], "test-key") - self.assertEqual(calls[0][1]["headers"]["xquik-api-contract"], API_CONTRACT) - self.assertEqual(calls[0][1]["params"]["limit"], 20) - self.assertFalse(calls[0][1]["allow_redirects"]) - self.assertEqual(len(collection["messages"]), 2) - self.assertEqual(collection["subject_candidates"], ["Example User", "example_user"]) - self.assertTrue(collection["metadata"]["has_more"]) - self.assertFalse(collection["metadata"]["pagination_followed"]) - self.assertNotIn("test-key", json.dumps(collection)) - - def test_collect_public_posts_enforces_limit_on_oversized_response(self) -> None: - collection = collect_public_posts( - api_key="test-key", - query="from:example_user", - limit=1, - query_type="Latest", - request_get=lambda *args, **kwargs: FakeResponse( - {"tweets": [sample_post("1"), sample_post("2")]} - ), - ) - - self.assertEqual([post["id"] for post in collection["messages"]], ["1"]) - - def test_collect_public_posts_rejects_missing_tweets(self) -> None: - with self.assertRaisesRegex(CollectorError, "missing the tweets list"): - collect_public_posts( - api_key="test-key", - query="from:example_user", - limit=20, - query_type="Latest", - request_get=lambda *args, **kwargs: FakeResponse({}), - ) - - def test_collect_public_posts_rejects_unknown_sort(self) -> None: - with self.assertRaisesRegex(CollectorError, "Sort must be Latest or Top"): - collect_public_posts( - api_key="test-key", - query="from:example_user", - limit=20, - query_type="Popular", - request_get=lambda *args, **kwargs: FakeResponse({"tweets": []}), - ) - - def test_collect_public_posts_sanitizes_authentication_error(self) -> None: - with self.assertRaisesRegex(CollectorError, "Authentication failed") as raised: - collect_public_posts( - api_key="secret-key", - query="from:example_user", - limit=20, - query_type="Latest", - request_get=lambda *args, **kwargs: FakeResponse({"secret": "body"}, 401), - ) - - self.assertNotIn("secret-key", str(raised.exception)) - self.assertNotIn("body", str(raised.exception)) - - def test_collect_public_posts_refuses_redirects_without_exposing_key(self) -> None: - with self.assertRaisesRegex(CollectorError, "Refusing to forward") as raised: - collect_public_posts( - api_key="secret-key", - query="from:example_user", - limit=20, - query_type="Latest", - request_get=lambda *args, **kwargs: FakeResponse({}, 302), - ) - - self.assertNotIn("secret-key", str(raised.exception)) - - def test_write_collection_requires_force_to_replace_output(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - output = Path(tmp_dir) / "candidates" / "x.json" - write_collection(output, {"messages": []}, force=False) - - with self.assertRaisesRegex(CollectorError, "Pass --force"): - write_collection(output, {"messages": [1]}, force=False) - - write_collection(output, {"messages": [1]}, force=True) - self.assertEqual(json.loads(output.read_text(encoding="utf-8")), {"messages": [1]}) - self.assertEqual([path.name for path in output.parent.iterdir()], ["x.json"]) - - @unittest.skipUnless(hasattr(os, "symlink"), "symbolic links are unavailable") - def test_write_collection_rejects_dangling_symlink(self) -> None: - with tempfile.TemporaryDirectory() as tmp_dir: - output = Path(tmp_dir) / "candidates" / "x.json" - target = Path(tmp_dir) / "outside.json" - output.parent.mkdir() - try: - output.symlink_to(target) - except OSError as error: - self.skipTest(f"cannot create symbolic links: {error}") - - with self.assertRaisesRegex(CollectorError, "symbolic link"): - write_collection(output, {"messages": []}, force=False) - with self.assertRaisesRegex(CollectorError, "symbolic link"): - write_collection(output, {"messages": []}, force=True) - - self.assertFalse(target.exists()) - - -if __name__ == "__main__": - unittest.main() diff --git a/tools/dingtalk_auto_collector.py b/tools/dingtalk_auto_collector.py deleted file mode 100644 index 4d451739..00000000 --- a/tools/dingtalk_auto_collector.py +++ /dev/null @@ -1,790 +0,0 @@ -#!/usr/bin/env python3 -""" -钉钉自动采集器 - -输入同事姓名,自动: - 1. 搜索钉钉用户,获取 userId - 2. 搜索他创建/编辑的文档和知识库内容 - 3. 拉取多维表格(如有) - 4. 消息记录(API 不支持历史拉取,自动切换浏览器方案) - 5. 输出统一格式,直接进入 Distilly 分析流程 - -钉钉限制说明: - 钉钉 Open API 不提供历史消息拉取接口, - 消息记录部分自动使用 Playwright 浏览器方案采集。 - -前置: - pip3 install requests playwright - playwright install chromium - python3 dingtalk_auto_collector.py --setup - -用法: - python3 dingtalk_auto_collector.py --name "张三" --output-dir ./knowledge/zhangsan - python3 dingtalk_auto_collector.py --name "张三" --skip-messages # 跳过消息采集 - python3 dingtalk_auto_collector.py --name "张三" --doc-limit 20 -""" - -from __future__ import annotations - -import json -import sys -import time -import argparse -import platform -from pathlib import Path -from datetime import datetime, timezone -from typing import Optional - -try: - import requests -except ImportError: - print("错误:请先安装依赖:pip3 install requests", file=sys.stderr) - sys.exit(1) - - -CONFIG_PATH = Path.home() / ".distilly" / "dingtalk_config.json" -LEGACY_CONFIG_PATH = Path.home() / ".colleague-skill" / "dingtalk_config.json" -API_BASE = "https://api.dingtalk.com" - - -# ─── 配置 ──────────────────────────────────────────────────────────────────── - -def load_config() -> dict: - config_path = CONFIG_PATH if CONFIG_PATH.exists() else LEGACY_CONFIG_PATH - if not config_path.exists(): - print("未找到配置,请先运行:python3 dingtalk_auto_collector.py --setup", file=sys.stderr) - sys.exit(1) - return json.loads(config_path.read_text(encoding="utf-8")) - - -def save_config(config: dict) -> None: - CONFIG_PATH.parent.mkdir(parents=True, exist_ok=True) - CONFIG_PATH.write_text(json.dumps(config, indent=2, ensure_ascii=False)) - CONFIG_PATH.chmod(0o600) - - -def setup_config() -> None: - print("=== 钉钉自动采集配置 ===\n") - print("请前往 https://open-dev.dingtalk.com 创建企业内部应用,开通以下权限:\n") - print(" 通讯录类:") - print(" qyapi_get_member_detail 查询用户详情") - print(" Contact.User.mobile 读取用户手机号(可选)") - print() - print(" 消息类(可选,仅用于发消息,历史消息需浏览器方案):") - print(" qyapi_robot_sendmsg 机器人发消息") - print() - print(" 文档类:") - print(" Doc.WorkSpace.READ 读取工作空间") - print(" Doc.File.READ 读取文件") - print() - print(" 多维表格:") - print(" Bitable.Record.READ 读取记录") - print() - - app_key = input("AppKey (ding_xxx): ").strip() - app_secret = input("AppSecret: ").strip() - - config = {"app_key": app_key, "app_secret": app_secret} - save_config(config) - print(f"\n✅ 配置已保存到 {CONFIG_PATH}") - print("\n注意:消息记录采集需要 Playwright,请确认已安装:") - print(" pip3 install playwright && playwright install chromium") - - -# ─── Token ─────────────────────────────────────────────────────────────────── - -_token_cache: dict = {} - - -def get_access_token(config: dict) -> str: - """获取钉钉 access_token,带缓存""" - now = time.time() - if _token_cache.get("token") and _token_cache.get("expire", 0) > now + 60: - return _token_cache["token"] - - resp = requests.post( - f"{API_BASE}/v1.0/oauth2/accessToken", - json={"appKey": config["app_key"], "appSecret": config["app_secret"]}, - timeout=10, - ) - data = resp.json() - - if "accessToken" not in data: - print(f"获取 token 失败:{data}", file=sys.stderr) - sys.exit(1) - - token = data["accessToken"] - _token_cache["token"] = token - _token_cache["expire"] = now + data.get("expireIn", 7200) - return token - - -def api_get(path: str, params: dict, config: dict) -> dict: - token = get_access_token(config) - resp = requests.get( - f"{API_BASE}{path}", - params=params, - headers={"x-acs-dingtalk-access-token": token}, - timeout=15, - ) - return resp.json() - - -def api_post(path: str, body: dict, config: dict) -> dict: - token = get_access_token(config) - resp = requests.post( - f"{API_BASE}{path}", - json=body, - headers={"x-acs-dingtalk-access-token": token}, - timeout=15, - ) - return resp.json() - - -# ─── 用户搜索 ───────────────────────────────────────────────────────────────── - -def find_user(name: str, config: dict) -> Optional[dict]: - """通过姓名搜索钉钉用户""" - print(f" 搜索用户:{name} ...", file=sys.stderr) - - data = api_post( - "/v1.0/contact/users/search", - {"searchText": name, "offset": 0, "size": 10}, - config, - ) - - users = data.get("list", []) or data.get("result", {}).get("list", []) - - if not users: - # 降级:通过部门遍历搜索 - print(" API 搜索无结果,尝试遍历通讯录 ...", file=sys.stderr) - users = search_users_by_dept(name, config) - - if not users: - print(f" 未找到用户:{name}", file=sys.stderr) - return None - - if len(users) == 1: - u = users[0] - print(f" 找到用户:{u.get('name')}({u.get('deptNameList', [''])[0] if isinstance(u.get('deptNameList'), list) else ''})", file=sys.stderr) - return u - - print(f"\n 找到 {len(users)} 个结果,请选择:") - for i, u in enumerate(users): - dept = u.get("deptNameList", [""]) - dept_str = dept[0] if isinstance(dept, list) and dept else "" - print(f" [{i+1}] {u.get('name')} {dept_str} {u.get('unionId', '')}") - - choice = input("\n 选择编号(默认 1):").strip() or "1" - try: - return users[int(choice) - 1] - except (ValueError, IndexError): - return users[0] - - -def search_users_by_dept(name: str, config: dict, dept_id: int = 1, depth: int = 0) -> list: - """递归遍历部门搜索用户(深度限制 3 层)""" - if depth > 3: - return [] - - results = [] - - # 获取部门用户列表 - data = api_post( - "/v1.0/contact/users/simplelist", - {"deptId": dept_id, "cursor": 0, "size": 100}, - config, - ) - users = data.get("list", []) - for u in users: - if name in u.get("name", ""): - # 获取详细信息 - detail = api_get(f"/v1.0/contact/users/{u.get('userId')}", {}, config) - results.append(detail.get("result", u)) - - # 获取子部门 - sub_data = api_get( - "/v1.0/contact/departments/listSubDepts", - {"deptId": dept_id}, - config, - ) - for sub in sub_data.get("result", []): - results.extend(search_users_by_dept(name, config, sub.get("deptId"), depth + 1)) - - return results - - -# ─── 文档采集 ───────────────────────────────────────────────────────────────── - -def list_workspaces(config: dict) -> list: - """获取所有工作空间""" - data = api_get("/v1.0/doc/workspaces", {"maxResults": 50}, config) - return data.get("workspaceModels", []) or data.get("result", {}).get("workspaceModels", []) - - -def search_docs_by_user(user_id: str, name: str, doc_limit: int, config: dict) -> list: - """搜索用户创建的文档""" - print(f" 搜索 {name} 的文档 ...", file=sys.stderr) - - # 方式一:全局搜索 - data = api_post( - "/v1.0/doc/search", - { - "keyword": name, - "size": doc_limit, - "offset": 0, - }, - config, - ) - - docs = [] - items = data.get("docList", []) or data.get("result", {}).get("docList", []) - - for item in items: - creator_id = item.get("creatorId", "") or item.get("creator", {}).get("userId", "") - # 过滤:只保留目标用户创建的 - if user_id and creator_id and creator_id != user_id: - continue - docs.append({ - "title": item.get("title", "无标题"), - "docId": item.get("docId", ""), - "spaceId": item.get("spaceId", ""), - "type": item.get("docType", ""), - "url": item.get("shareUrl", ""), - "creator": item.get("creatorName", name), - }) - - if not docs: - # 方式二:遍历工作空间找文档 - print(" 搜索无结果,遍历工作空间 ...", file=sys.stderr) - workspaces = list_workspaces(config) - for ws in workspaces[:5]: # 最多查 5 个空间 - ws_id = ws.get("spaceId") or ws.get("workspaceId") - if not ws_id: - continue - files_data = api_get( - f"/v1.0/doc/workspaces/{ws_id}/files", - {"maxResults": 20, "orderBy": "modified_time", "order": "DESC"}, - config, - ) - for f in files_data.get("files", []): - creator_id = f.get("creatorId", "") - if user_id and creator_id and creator_id != user_id: - continue - docs.append({ - "title": f.get("fileName", "无标题"), - "docId": f.get("docId", ""), - "spaceId": ws_id, - "type": f.get("docType", ""), - "url": f.get("shareUrl", ""), - "creator": name, - }) - - print(f" 找到 {len(docs)} 篇文档", file=sys.stderr) - return docs[:doc_limit] - - -def fetch_doc_content(doc_id: str, space_id: str, config: dict) -> str: - """拉取单篇文档的文本内容""" - # 方式一:直接获取文档内容 - data = api_get( - f"/v1.0/doc/workspaces/{space_id}/files/{doc_id}/content", - {}, - config, - ) - - content = ( - data.get("content") - or data.get("result", {}).get("content") - or data.get("markdown") - or data.get("result", {}).get("markdown") - or "" - ) - - if content: - return content - - # 方式二:获取下载链接后下载 - dl_data = api_get( - f"/v1.0/doc/workspaces/{space_id}/files/{doc_id}/download", - {}, - config, - ) - dl_url = dl_data.get("downloadUrl") or dl_data.get("result", {}).get("downloadUrl") - if dl_url: - try: - resp = requests.get(dl_url, timeout=15) - return resp.text - except Exception: - pass - - return "" - - -def collect_docs(user: dict, doc_limit: int, config: dict) -> str: - """采集目标用户的文档""" - user_id = user.get("userId", "") - name = user.get("name", "") - - docs = search_docs_by_user(user_id, name, doc_limit, config) - if not docs: - return f"# 文档内容\n\n未找到 {name} 相关文档\n" - - lines = [ - "# 文档内容(钉钉自动采集)", - f"目标:{name}", - f"共 {len(docs)} 篇", - "", - ] - - for doc in docs: - title = doc.get("title", "无标题") - doc_id = doc.get("docId", "") - space_id = doc.get("spaceId", "") - url = doc.get("url", "") - - if not doc_id or not space_id: - continue - - print(f" 拉取文档:{title} ...", file=sys.stderr) - content = fetch_doc_content(doc_id, space_id, config) - - if not content or len(content.strip()) < 20: - print(f" 内容为空,跳过", file=sys.stderr) - continue - - lines += [ - "---", - f"## 《{title}》", - f"链接:{url}", - f"创建人:{doc.get('creator', '')}", - "", - content.strip(), - "", - ] - - return "\n".join(lines) - - -# ─── 多维表格 ───────────────────────────────────────────────────────────────── - -def search_bitables(user_id: str, name: str, config: dict) -> list: - """搜索目标用户的多维表格""" - print(f" 搜索 {name} 的多维表格 ...", file=sys.stderr) - - data = api_post( - "/v1.0/doc/search", - {"keyword": name, "size": 20, "offset": 0, "docTypes": ["bitable"]}, - config, - ) - - tables = [] - for item in data.get("docList", []): - if item.get("docType") != "bitable": - continue - creator_id = item.get("creatorId", "") - if user_id and creator_id and creator_id != user_id: - continue - tables.append(item) - - print(f" 找到 {len(tables)} 个多维表格", file=sys.stderr) - return tables - - -def fetch_bitable_content(base_id: str, config: dict) -> str: - """拉取多维表格内容""" - # 获取所有 sheet - sheets_data = api_get( - f"/v1.0/bitable/bases/{base_id}/sheets", - {}, - config, - ) - sheets = sheets_data.get("sheets", []) or sheets_data.get("result", {}).get("sheets", []) - - if not sheets: - return "(多维表格为空或无权限)\n" - - lines = [] - for sheet in sheets: - sheet_id = sheet.get("sheetId") or sheet.get("id") - sheet_name = sheet.get("name", sheet_id) - - # 获取字段 - fields_data = api_get( - f"/v1.0/bitable/bases/{base_id}/sheets/{sheet_id}/fields", - {"maxResults": 100}, - config, - ) - fields = [f.get("name", "") for f in fields_data.get("fields", [])] - - # 获取记录 - records_data = api_get( - f"/v1.0/bitable/bases/{base_id}/sheets/{sheet_id}/records", - {"maxResults": 200}, - config, - ) - records = records_data.get("records", []) or records_data.get("result", {}).get("records", []) - - lines.append(f"### 表:{sheet_name}") - lines.append("") - - if fields: - lines.append("| " + " | ".join(fields) + " |") - lines.append("| " + " | ".join(["---"] * len(fields)) + " |") - - for rec in records: - row_data = rec.get("fields", {}) - row = [] - for f in fields: - val = row_data.get(f, "") - if isinstance(val, list): - val = " ".join( - v.get("text", str(v)) if isinstance(v, dict) else str(v) - for v in val - ) - row.append(str(val).replace("|", "|").replace("\n", " ")) - lines.append("| " + " | ".join(row) + " |") - - lines.append("") - - return "\n".join(lines) - - -def collect_bitables(user: dict, config: dict) -> str: - """采集目标用户的多维表格""" - user_id = user.get("userId", "") - name = user.get("name", "") - - tables = search_bitables(user_id, name, config) - if not tables: - return f"# 多维表格\n\n未找到 {name} 的多维表格\n" - - lines = [ - "# 多维表格(钉钉自动采集)", - f"目标:{name}", - f"共 {len(tables)} 个", - "", - ] - - for t in tables: - title = t.get("title", "无标题") - doc_id = t.get("docId", "") - print(f" 拉取多维表格:{title} ...", file=sys.stderr) - - content = fetch_bitable_content(doc_id, config) - lines += [ - "---", - f"## 《{title}》", - "", - content, - ] - - return "\n".join(lines) - - -# ─── 消息记录(浏览器方案)──────────────────────────────────────────────────── - -def get_default_chrome_profile() -> str: - system = platform.system() - if system == "Darwin": - return str(Path.home() / "Library/Application Support/Google/Chrome/Default") - elif system == "Linux": - return str(Path.home() / ".config/google-chrome/Default") - elif system == "Windows": - import os - return str(Path(os.environ.get("LOCALAPPDATA", "")) / "Google/Chrome/User Data/Default") - return str(Path.home() / ".config/google-chrome/Default") - - -def collect_messages_browser( - name: str, - msg_limit: int, - chrome_profile: Optional[str], - headless: bool, -) -> str: - """通过 Playwright 浏览器抓取钉钉网页版消息记录""" - try: - from playwright.sync_api import sync_playwright - except ImportError: - return ( - "# 消息记录\n\n" - "⚠️ 未安装 Playwright,无法采集消息记录。\n" - "请运行:pip3 install playwright && playwright install chromium\n" - ) - - import re - - profile = chrome_profile or get_default_chrome_profile() - print(f" 启动浏览器抓取钉钉消息({'无头' if headless else '有界面'})...", file=sys.stderr) - - messages = [] - - with sync_playwright() as p: - try: - ctx = p.chromium.launch_persistent_context( - user_data_dir=profile, - headless=headless, - args=["--disable-blink-features=AutomationControlled"], - ignore_default_args=["--enable-automation"], - viewport={"width": 1280, "height": 900}, - ) - except Exception as e: - return f"# 消息记录\n\n⚠️ 无法启动浏览器:{e}\n" - - page = ctx.new_page() - - # 打开钉钉网页版 - page.goto("https://im.dingtalk.com", wait_until="domcontentloaded", timeout=20000) - time.sleep(3) - - # 检查登录状态 - if "login" in page.url.lower() or page.query_selector(".login-wrap"): - if headless: - ctx.close() - return ( - "# 消息记录\n\n" - "⚠️ 检测到未登录。请用 --show-browser 参数重新运行,在弹出窗口中登录钉钉。\n" - ) - print(" 请在浏览器中登录钉钉,登录完成后按回车继续...", file=sys.stderr) - input() - - # 搜索目标联系人的消息 - try: - # 点击搜索框 - search_selectors = [ - '[placeholder*="搜索"]', - '.search-input', - '[data-testid="search"]', - '.im-search', - ] - for sel in search_selectors: - el = page.query_selector(sel) - if el: - el.click() - time.sleep(0.5) - page.keyboard.type(name) - time.sleep(2) - break - - # 点击第一个结果 - result_selectors = [ - '.search-result-item', - '.contact-item', - '.result-item', - ] - for sel in result_selectors: - result = page.query_selector(sel) - if result: - result.click() - time.sleep(2) - break - except Exception as e: - print(f" 自动导航失败:{e}", file=sys.stderr) - if not headless: - print(f" 请手动打开与「{name}」的对话,然后按回车继续...", file=sys.stderr) - input() - - # 向上滚动加载历史消息 - print(" 加载历史消息 ...", file=sys.stderr) - for _ in range(15): - page.keyboard.press("Control+Home") - time.sleep(1) - page.evaluate("window.scrollTo(0, 0)") - time.sleep(0.8) - - time.sleep(2) - - # 提取消息 - raw_messages = page.evaluate(f""" - () => {{ - const target = "{name}"; - const results = []; - const selectors = [ - '.message-item-content-container', - '.im-message-item', - '[data-message-id]', - '.msg-wrap', - ]; - - let items = []; - for (const sel of selectors) {{ - items = document.querySelectorAll(sel); - if (items.length > 0) break; - }} - - items.forEach(item => {{ - const senderEl = item.querySelector('.sender-name, .nick-name, .name'); - const contentEl = item.querySelector( - '.message-text, .text-content, .msg-content, .im-richtext' - ); - const timeEl = item.querySelector('.message-time, .time, .msg-time'); - - const sender = senderEl ? senderEl.innerText.trim() : ''; - const content = contentEl ? contentEl.innerText.trim() : ''; - const time = timeEl ? timeEl.innerText.trim() : ''; - - if (!content) return; - if (target && !sender.includes(target)) return; - if (['[图片]','[文件]','[表情]','[语音]'].includes(content)) return; - - results.push({{ sender, content, time }}); - }}); - - return results.slice(-{msg_limit}); - }} - """) - - ctx.close() - messages = raw_messages or [] - - if not messages: - return ( - "# 消息记录\n\n" - f"⚠️ 未能自动提取 {name} 的消息。\n" - "可能原因:钉钉网页版 DOM 结构变化,或未找到对话。\n" - "建议手动截图聊天记录后上传。\n" - ) - - long_msgs = [m for m in messages if len(m.get("content", "")) > 50] - short_msgs = [m for m in messages if len(m.get("content", "")) <= 50] - - lines = [ - "# 消息记录(钉钉浏览器采集)", - f"目标:{name}", - f"共 {len(messages)} 条", - "注意:钉钉 API 不支持历史消息拉取,本内容通过浏览器采集", - "", - "---", - "", - "## 长消息(观点/决策/技术类)", - "", - ] - for m in long_msgs: - lines.append(f"[{m.get('time', '')}] {m.get('content', '')}") - lines.append("") - - lines += ["---", "", "## 日常消息(风格参考)", ""] - for m in short_msgs[:300]: - lines.append(f"[{m.get('time', '')}] {m.get('content', '')}") - - return "\n".join(lines) - - -# ─── 主流程 ─────────────────────────────────────────────────────────────────── - -def collect_all( - name: str, - output_dir: Path, - msg_limit: int, - doc_limit: int, - skip_messages: bool, - chrome_profile: Optional[str], - headless: bool, - config: dict, -) -> dict: - output_dir.mkdir(parents=True, exist_ok=True) - results = {} - - print(f"\n🔍 开始采集(钉钉):{name}\n", file=sys.stderr) - - # Step 1: 搜索用户 - user = find_user(name, config) - if not user: - print(f"❌ 未找到用户:{name}", file=sys.stderr) - sys.exit(1) - - print(f" 用户 ID:{user.get('userId', '')} 部门:{user.get('deptNameList', [''])[0] if isinstance(user.get('deptNameList'), list) and user.get('deptNameList') else ''}", file=sys.stderr) - - # Step 2: 文档 - print(f"\n📄 采集文档(上限 {doc_limit} 篇)...", file=sys.stderr) - try: - doc_content = collect_docs(user, doc_limit, config) - doc_path = output_dir / "docs.txt" - doc_path.write_text(doc_content, encoding="utf-8") - results["docs"] = str(doc_path) - print(f" ✅ 文档 → {doc_path}", file=sys.stderr) - except Exception as e: - print(f" ⚠️ 文档采集失败:{e}", file=sys.stderr) - - # Step 3: 多维表格 - print(f"\n📊 采集多维表格 ...", file=sys.stderr) - try: - bitable_content = collect_bitables(user, config) - bt_path = output_dir / "bitables.txt" - bt_path.write_text(bitable_content, encoding="utf-8") - results["bitables"] = str(bt_path) - print(f" ✅ 多维表格 → {bt_path}", file=sys.stderr) - except Exception as e: - print(f" ⚠️ 多维表格采集失败:{e}", file=sys.stderr) - - # Step 4: 消息记录(浏览器方案) - if not skip_messages: - print(f"\n📨 采集消息记录(浏览器方案,上限 {msg_limit} 条)...", file=sys.stderr) - print(f" ℹ️ 钉钉 API 不支持历史消息拉取,自动切换浏览器方案", file=sys.stderr) - try: - msg_content = collect_messages_browser(name, msg_limit, chrome_profile, headless) - msg_path = output_dir / "messages.txt" - msg_path.write_text(msg_content, encoding="utf-8") - results["messages"] = str(msg_path) - print(f" ✅ 消息记录 → {msg_path}", file=sys.stderr) - except Exception as e: - print(f" ⚠️ 消息采集失败:{e}", file=sys.stderr) - else: - print(f"\n📨 跳过消息采集(--skip-messages)", file=sys.stderr) - - # 写摘要 - summary = { - "name": name, - "user_id": user.get("userId", ""), - "platform": "dingtalk", - "department": user.get("deptNameList", []), - "collected_at": datetime.now(timezone.utc).isoformat(), - "files": results, - "notes": "消息记录通过浏览器采集,钉钉 API 不支持历史消息拉取", - } - (output_dir / "collection_summary.json").write_text( - json.dumps(summary, ensure_ascii=False, indent=2) - ) - - print(f"\n✅ 采集完成 → {output_dir}", file=sys.stderr) - print(f" 文件:{', '.join(results.keys())}", file=sys.stderr) - return results - - -def main() -> None: - parser = argparse.ArgumentParser(description="钉钉数据自动采集器") - parser.add_argument("--setup", action="store_true", help="初始化配置") - parser.add_argument("--name", help="同事姓名") - parser.add_argument("--output-dir", default=None, help="输出目录") - parser.add_argument("--msg-limit", type=int, default=500, help="最多采集消息条数(默认 500)") - parser.add_argument("--doc-limit", type=int, default=20, help="最多采集文档篇数(默认 20)") - parser.add_argument("--skip-messages", action="store_true", help="跳过消息记录采集") - parser.add_argument("--chrome-profile", default=None, help="Chrome Profile 路径") - parser.add_argument("--show-browser", action="store_true", help="显示浏览器窗口(调试/首次登录)") - - args = parser.parse_args() - - if args.setup: - setup_config() - return - - if not args.name: - parser.error("请提供 --name") - - config = load_config() - output_dir = Path(args.output_dir) if args.output_dir else Path(f"./knowledge/{args.name}") - - collect_all( - name=args.name, - output_dir=output_dir, - msg_limit=args.msg_limit, - doc_limit=args.doc_limit, - skip_messages=args.skip_messages, - chrome_profile=args.chrome_profile, - headless=not args.show_browser, - config=config, - ) - - -if __name__ == "__main__": - main() diff --git a/tools/email_parser.py b/tools/email_parser.py deleted file mode 100644 index fb5e1fa3..00000000 --- a/tools/email_parser.py +++ /dev/null @@ -1,339 +0,0 @@ -#!/usr/bin/env python3 -""" -邮件解析器 - -支持格式: -1. .eml 文件(标准邮件格式) -2. .txt 文件(纯文本邮件记录) -3. .mbox 文件(多封邮件合集) - -用法: - python email_parser.py --file emails.eml --target "zhangsan@company.com" --output output.txt - python email_parser.py --file inbox.mbox --target "张三" --output output.txt -""" - -import email -import email.policy -import mailbox -import re -import sys -import argparse -from pathlib import Path -from email.header import decode_header -from html.parser import HTMLParser - - -class HTMLTextExtractor(HTMLParser): - """从 HTML 邮件内容中提取纯文本""" - - def __init__(self): - super().__init__() - self.result = [] - self._skip = False - - def handle_starttag(self, tag, attrs): - if tag in ("script", "style"): - self._skip = True - - def handle_endtag(self, tag): - if tag in ("script", "style"): - self._skip = False - if tag in ("p", "br", "div", "tr"): - self.result.append("\n") - - def handle_data(self, data): - if not self._skip: - self.result.append(data) - - def get_text(self): - return re.sub(r"\n{3,}", "\n\n", "".join(self.result)).strip() - - -def decode_mime_str(s: str) -> str: - """解码 MIME 编码的邮件头字段""" - if not s: - return "" - parts = decode_header(s) - result = [] - for part, charset in parts: - if isinstance(part, bytes): - charset = charset or "utf-8" - try: - result.append(part.decode(charset, errors="replace")) - except Exception: - result.append(part.decode("utf-8", errors="replace")) - else: - result.append(str(part)) - return "".join(result) - - -def extract_email_body(msg) -> str: - """从邮件对象中提取正文文本""" - body = "" - - if msg.is_multipart(): - for part in msg.walk(): - content_type = part.get_content_type() - disposition = str(part.get("Content-Disposition", "")) - - if "attachment" in disposition: - continue - - if content_type == "text/plain": - payload = part.get_payload(decode=True) - charset = part.get_content_charset() or "utf-8" - try: - body = payload.decode(charset, errors="replace") - break - except Exception: - body = payload.decode("utf-8", errors="replace") - break - - elif content_type == "text/html" and not body: - payload = part.get_payload(decode=True) - charset = part.get_content_charset() or "utf-8" - try: - html = payload.decode(charset, errors="replace") - except Exception: - html = payload.decode("utf-8", errors="replace") - extractor = HTMLTextExtractor() - extractor.feed(html) - body = extractor.get_text() - else: - payload = msg.get_payload(decode=True) - if payload: - charset = msg.get_content_charset() or "utf-8" - try: - body = payload.decode(charset, errors="replace") - except Exception: - body = payload.decode("utf-8", errors="replace") - - # 清理引用内容(Re: 时的原文引用) - body = re.sub(r"\n>.*", "", body) - body = re.sub(r"\n-{3,}.*?原始邮件.*?\n", "\n", body, flags=re.DOTALL) - body = re.sub(r"\n_{3,}\n.*", "", body, flags=re.DOTALL) - - return body.strip() - - -def is_from_target(from_field: str, target: str) -> bool: - """判断邮件是否来自目标人""" - from_str = decode_mime_str(from_field).lower() - target_lower = target.lower() - return target_lower in from_str - - -def parse_eml_file(file_path: str, target: str) -> list[dict]: - """解析单个 .eml 文件""" - with open(file_path, "rb") as f: - msg = email.message_from_binary_file(f, policy=email.policy.default) - - from_field = str(msg.get("From", "")) - if not is_from_target(from_field, target): - return [] - - subject = decode_mime_str(str(msg.get("Subject", ""))) - date = str(msg.get("Date", "")) - body = extract_email_body(msg) - - if not body: - return [] - - return [{ - "from": decode_mime_str(from_field), - "subject": subject, - "date": date, - "body": body, - }] - - -def parse_mbox_file(file_path: str, target: str) -> list[dict]: - """解析 .mbox 文件(多封邮件合集)""" - results = [] - mbox = mailbox.mbox(file_path) - - for msg in mbox: - from_field = str(msg.get("From", "")) - if not is_from_target(from_field, target): - continue - - subject = decode_mime_str(str(msg.get("Subject", ""))) - date = str(msg.get("Date", "")) - body = extract_email_body(msg) - - if not body: - continue - - results.append({ - "from": decode_mime_str(from_field), - "subject": subject, - "date": date, - "body": body, - }) - - return results - - -def parse_txt_file(file_path: str, target: str) -> list[dict]: - """ - 解析纯文本格式的邮件记录 - 支持简单的分隔格式: - From: xxx - Subject: xxx - Date: xxx - --- - 正文内容 - === - """ - results = [] - - with open(file_path, "r", encoding="utf-8") as f: - content = f.read() - - # 尝试按分隔符切割多封邮件 - emails_raw = re.split(r"\n={3,}\n|\n-{3,}\n(?=From:)", content) - - for raw in emails_raw: - from_match = re.search(r"^From:\s*(.+)$", raw, re.MULTILINE) - subject_match = re.search(r"^Subject:\s*(.+)$", raw, re.MULTILINE) - date_match = re.search(r"^Date:\s*(.+)$", raw, re.MULTILINE) - - from_field = from_match.group(1).strip() if from_match else "" - if not is_from_target(from_field, target): - continue - - # 提取正文(去掉头部字段后的内容) - body = re.sub(r"^(From|To|Subject|Date|CC|BCC):.*\n?", "", raw, flags=re.MULTILINE) - body = body.strip() - - if not body: - continue - - results.append({ - "from": from_field, - "subject": subject_match.group(1).strip() if subject_match else "", - "date": date_match.group(1).strip() if date_match else "", - "body": body, - }) - - return results - - -def classify_emails(emails: list[dict]) -> dict: - """ - 对邮件按内容分类: - - 长邮件(正文 > 200 字):技术方案、观点陈述 - - 决策类:包含明确判断的邮件 - - 日常沟通:短邮件 - """ - long_emails = [] - decision_emails = [] - daily_emails = [] - - decision_keywords = [ - "同意", "不同意", "建议", "方案", "觉得", "应该", "决定", "确认", - "approve", "reject", "lgtm", "suggest", "recommend", "think", - "我的看法", "我认为", "我觉得", "需要", "必须", "不需要" - ] - - for e in emails: - body = e["body"] - - if len(body) > 200: - long_emails.append(e) - elif any(kw in body.lower() for kw in decision_keywords): - decision_emails.append(e) - else: - daily_emails.append(e) - - return { - "long_emails": long_emails, - "decision_emails": decision_emails, - "daily_emails": daily_emails, - "total_count": len(emails), - } - - -def format_output(target: str, classified: dict) -> str: - """格式化输出,供 AI 分析使用""" - lines = [ - f"# 邮件提取结果", - f"目标人物:{target}", - f"总邮件数:{classified['total_count']}", - "", - "---", - "", - "## 长邮件(技术方案/观点类,权重最高)", - "", - ] - - for e in classified["long_emails"]: - lines.append(f"**主题:{e['subject']}** [{e['date']}]") - lines.append(e["body"]) - lines.append("") - lines.append("---") - lines.append("") - - lines += [ - "## 决策类邮件", - "", - ] - - for e in classified["decision_emails"]: - lines.append(f"**主题:{e['subject']}** [{e['date']}]") - lines.append(e["body"]) - lines.append("") - - lines += [ - "---", - "", - "## 日常沟通(风格参考)", - "", - ] - - for e in classified["daily_emails"][:30]: - lines.append(f"**{e['subject']}**:{e['body'][:200]}") - lines.append("") - - return "\n".join(lines) - - -def main(): - parser = argparse.ArgumentParser(description="解析邮件文件,提取目标人发出的邮件") - parser.add_argument("--file", required=True, help="输入文件路径(.eml / .mbox / .txt)") - parser.add_argument("--target", required=True, help="目标人物(邮箱地址或姓名)") - parser.add_argument("--output", default=None, help="输出文件路径(默认打印到 stdout)") - - args = parser.parse_args() - - file_path = Path(args.file) - if not file_path.exists(): - print(f"错误:文件不存在 {file_path}", file=sys.stderr) - sys.exit(1) - - suffix = file_path.suffix.lower() - - if suffix == ".eml": - emails = parse_eml_file(str(file_path), args.target) - elif suffix == ".mbox": - emails = parse_mbox_file(str(file_path), args.target) - else: - emails = parse_txt_file(str(file_path), args.target) - - if not emails: - print(f"警告:未找到来自 '{args.target}' 的邮件", file=sys.stderr) - print("提示:请检查目标名称/邮箱是否与文件中的 From 字段一致", file=sys.stderr) - - classified = classify_emails(emails) - output = format_output(args.target, classified) - - if args.output: - with open(args.output, "w", encoding="utf-8") as f: - f.write(output) - print(f"已输出到 {args.output},共 {len(emails)} 封邮件") - else: - print(output) - - -if __name__ == "__main__": - main() diff --git a/tools/feishu_auto_collector.py b/tools/feishu_auto_collector.py deleted file mode 100644 index 19739e3f..00000000 --- a/tools/feishu_auto_collector.py +++ /dev/null @@ -1,960 +0,0 @@ -#!/usr/bin/env python3 -""" -飞书自动采集器 - -输入同事姓名,自动: - 1. 搜索飞书用户,获取 user_id - 2. 找到与他共同的群聊,拉取他的消息记录 - 3. 拉取私聊消息(需要 user_access_token) - 4. 搜索他创建/编辑的文档和 Wiki - 5. 拉取文档内容 - 6. 拉取多维表格(如有) - 7. 输出统一格式,直接进入 Distilly 分析流程 - -前置: - python3 feishu_auto_collector.py --setup # 配置 App ID / Secret(一次性) - -私聊采集(需额外步骤): - 1. 飞书应用开通用户权限:im:message, im:chat - 2. 获取 OAuth 授权码: - 浏览器打开: https://open.feishu.cn/open-apis/authen/v1/authorize?app_id={APP_ID}&redirect_uri=http://www.example.com&scope=im:message%20im:chat - 授权后从地址栏复制 code - 3. 换取 token: - python3 feishu_auto_collector.py --exchange-code {CODE} - 4. 采集时指定私聊 chat_id: - python3 feishu_auto_collector.py --name "张三" --p2p-chat-id oc_xxx - -用法: - # 群聊采集(原有方式) - python3 feishu_auto_collector.py --name "张三" --output-dir ./knowledge/zhangsan - python3 feishu_auto_collector.py --name "张三" --msg-limit 1000 --doc-limit 20 - - # 私聊采集 - python3 feishu_auto_collector.py --name "张三" --p2p-chat-id oc_xxx - - # 直接指定 open_id + 私聊(跳过用户搜索) - python3 feishu_auto_collector.py --open-id ou_xxx --p2p-chat-id oc_xxx --name "张三" - - # 换取 user_access_token - python3 feishu_auto_collector.py --exchange-code {CODE} -""" - -from __future__ import annotations - -import json -import sys -import time -import argparse -from pathlib import Path -from datetime import datetime, timezone -from typing import Optional - -try: - import requests -except ImportError: - print("错误:请先安装 requests:pip3 install requests", file=sys.stderr) - sys.exit(1) - - -CONFIG_PATH = Path.home() / ".distilly" / "feishu_config.json" -LEGACY_CONFIG_PATH = Path.home() / ".colleague-skill" / "feishu_config.json" -BASE_URL = "https://open.feishu.cn/open-apis" - - -# ─── 配置 ──────────────────────────────────────────────────────────────────── - -def load_config() -> dict: - config_path = CONFIG_PATH if CONFIG_PATH.exists() else LEGACY_CONFIG_PATH - if not config_path.exists(): - print("未找到配置,请先运行:python3 feishu_auto_collector.py --setup", file=sys.stderr) - sys.exit(1) - return json.loads(config_path.read_text()) - - -def save_config(config: dict) -> None: - CONFIG_PATH.parent.mkdir(parents=True, exist_ok=True) - CONFIG_PATH.write_text(json.dumps(config, indent=2, ensure_ascii=False)) - CONFIG_PATH.chmod(0o600) - - -def setup_config() -> None: - print("=== 飞书自动采集配置 ===\n") - print("请前往 https://open.feishu.cn 创建企业自建应用,开通以下权限:") - print() - print(" 消息类(应用权限,用于群聊采集):") - print(" im:message:readonly 读取消息") - print(" im:chat:readonly 读取群聊信息") - print(" im:chat.members:readonly 读取群成员") - print() - print(" 消息类(用户权限,用于私聊采集):") - print(" im:message 以用户身份读取/发送消息") - print(" im:chat 以用户身份读取会话列表") - print() - print(" 用户类:") - print(" contact:user.base:readonly 读取用户基本信息") - print(" contact:department.base:readonly 遍历部门查找用户(按姓名搜索必需)") - print() - print(" 文档类:") - print(" docs:doc:readonly 读取文档") - print(" wiki:wiki:readonly 读取知识库") - print(" drive:drive:readonly 搜索云盘文件") - print() - print(" 多维表格:") - print(" bitable:app:readonly 读取多维表格") - print() - print(" ─── 私聊采集说明 ───") - print(" 私聊消息必须通过 user_access_token 获取(应用身份无权访问私聊)。") - print(" 获取方式:OAuth 授权,授权链接格式:") - print(" https://open.feishu.cn/open-apis/authen/v1/authorize?app_id={APP_ID}&redirect_uri={REDIRECT}&scope=im:message%20im:chat") - print(" 授权后从回调 URL 中取 code,用 --exchange-code 换取 token。") - print() - - app_id = input("App ID (cli_xxx): ").strip() - app_secret = input("App Secret: ").strip() - - config = {"app_id": app_id, "app_secret": app_secret} - - print("\n是否配置 user_access_token?(用于私聊消息采集,可跳过)") - user_token = input("user_access_token (留空跳过): ").strip() - if user_token: - config["user_access_token"] = user_token - p2p_chat_id = input("私聊 chat_id (留空跳过): ").strip() - if p2p_chat_id: - config["p2p_chat_id"] = p2p_chat_id - - save_config(config) - print(f"\n✅ 配置已保存到 {CONFIG_PATH}") - - -# ─── Token ─────────────────────────────────────────────────────────────────── - -_token_cache: dict = {} - - -def get_tenant_token(config: dict) -> str: - """获取 tenant_access_token,带缓存(有效期约 2 小时)""" - now = time.time() - if _token_cache.get("token") and _token_cache.get("expire", 0) > now + 60: - return _token_cache["token"] - - resp = requests.post( - f"{BASE_URL}/auth/v3/tenant_access_token/internal", - json={"app_id": config["app_id"], "app_secret": config["app_secret"]}, - timeout=10, - ) - data = resp.json() - if data.get("code") != 0: - print(f"获取 token 失败:{data}", file=sys.stderr) - sys.exit(1) - - token = data["tenant_access_token"] - _token_cache["token"] = token - _token_cache["expire"] = now + data.get("expire", 7200) - return token - - -def api_get(path: str, params: dict, config: dict, use_user_token: bool = False) -> dict: - if use_user_token and config.get("user_access_token"): - token = config["user_access_token"] - else: - token = get_tenant_token(config) - resp = requests.get( - f"{BASE_URL}{path}", - params=params, - headers={"Authorization": f"Bearer {token}"}, - timeout=15, - ) - return resp.json() - - -def api_post(path: str, body: dict, config: dict, use_user_token: bool = False) -> dict: - if use_user_token and config.get("user_access_token"): - token = config["user_access_token"] - else: - token = get_tenant_token(config) - resp = requests.post( - f"{BASE_URL}{path}", - json=body, - headers={"Authorization": f"Bearer {token}"}, - timeout=15, - ) - return resp.json() - - -def exchange_code_for_token(code: str, config: dict) -> dict: - """用 OAuth 授权码换取 user_access_token""" - app_token = get_tenant_token(config) - resp = requests.post( - f"{BASE_URL}/authen/v1/oidc/access_token", - headers={"Authorization": f"Bearer {app_token}"}, - json={"grant_type": "authorization_code", "code": code}, - timeout=10, - ) - data = resp.json() - if data.get("code") != 0: - print(f"换取 token 失败:{data}", file=sys.stderr) - return {} - return data.get("data", {}) - - -# ─── 用户搜索 ───────────────────────────────────────────────────────────────── - -def _find_user_by_contact(name: str, config: dict) -> Optional[dict]: - """通过邮箱或手机号查找用户(使用 tenant_access_token)""" - # 判断输入类型 - emails, mobiles = [], [] - if "@" in name: - emails = [name] - elif name.replace("+", "").replace("-", "").isdigit(): - mobiles = [name] - else: - return None # 不是邮箱或手机号,跳过 - - body = {} - if emails: - body["emails"] = emails - if mobiles: - body["mobiles"] = mobiles - - data = api_post("/contact/v3/users/batch_get_id", body, config) - if data.get("code") != 0: - print(f" 邮箱/手机号查找失败(code={data.get('code')}):{data.get('msg')}", file=sys.stderr) - return None - - user_list = data.get("data", {}).get("user_list", []) - for item in user_list: - user_id = item.get("user_id") - if user_id: - # 获取用户详情 - detail = api_get(f"/contact/v3/users/{user_id}", {"user_id_type": "user_id"}, config) - if detail.get("code") == 0: - user_data = detail.get("data", {}).get("user", {}) - print(f" 找到用户:{user_data.get('name', user_id)}", file=sys.stderr) - return user_data - # 如果详情拉不到,返回基本信息 - return {"user_id": user_id, "open_id": item.get("open_id", ""), "name": name} - - return None - - -def _find_user_by_department(name: str, config: dict) -> Optional[dict]: - """遍历部门查找用户(使用 tenant_access_token,需要 contact:department.base:readonly)""" - print(f" 通过部门遍历查找 {name} ...", file=sys.stderr) - - # 递归获取所有部门 ID - dept_ids = ["0"] # 0 = 根部门 - queue = ["0"] - while queue: - parent_id = queue.pop(0) - data = api_get( - f"/contact/v3/departments/{parent_id}/children", - {"page_size": 50, "fetch_child": False}, - config, - ) - if data.get("code") != 0: - if parent_id == "0": - print(f" 部门遍历失败(code={data.get('code')}):{data.get('msg')}", file=sys.stderr) - print(f" 请确认已开通 contact:department.base:readonly 权限", file=sys.stderr) - return None - continue - - children = data.get("data", {}).get("items", []) - for child in children: - child_id = child.get("department_id", "") - if child_id: - dept_ids.append(child_id) - queue.append(child_id) - - print(f" 共 {len(dept_ids)} 个部门,搜索用户 ...", file=sys.stderr) - - # 在每个部门中查找用户 - matches = [] - for dept_id in dept_ids: - page_token = None - while True: - params = {"department_id": dept_id, "page_size": 50} - if page_token: - params["page_token"] = page_token - - data = api_get("/contact/v3/users/find_by_department", params, config) - if data.get("code") != 0: - break - - users = data.get("data", {}).get("items", []) - for u in users: - uname = u.get("name", "") - en_name = u.get("en_name", "") - if name in uname or name in en_name or uname == name or en_name == name: - matches.append(u) - - if not data.get("data", {}).get("has_more"): - break - page_token = data.get("data", {}).get("page_token") - - if len(matches) >= 10: - break # 够了 - - return _select_user(matches, name) - - -def _select_user(users: list, name: str) -> Optional[dict]: - """从候选列表中选择用户""" - if not users: - print(f" 未找到用户:{name}", file=sys.stderr) - return None - - # 去重(按 user_id) - seen = set() - deduped = [] - for u in users: - uid = u.get("user_id", u.get("open_id", id(u))) - if uid not in seen: - seen.add(uid) - deduped.append(u) - users = deduped - - if len(users) == 1: - u = users[0] - dept_ids = u.get("department_ids", []) - print(f" 找到用户:{u.get('name')}(部门:{dept_ids[0] if dept_ids else ''})", file=sys.stderr) - return u - - # 多个结果,让用户选择 - print(f"\n 找到 {len(users)} 个结果,请选择:") - for i, u in enumerate(users): - dept_ids = u.get("department_ids", []) - dept_str = dept_ids[0] if dept_ids else "" - en = u.get("en_name", "") - label = f"{u.get('name', '')} ({en})" if en else u.get("name", "") - print(f" [{i+1}] {label} dept={dept_str} uid={u.get('user_id', '')}") - - choice = input("\n 选择编号(默认 1):").strip() or "1" - try: - idx = int(choice) - 1 - return users[idx] - except (ValueError, IndexError): - return users[0] - - -def find_user(name: str, config: dict) -> Optional[dict]: - """搜索飞书用户 - - 策略: - 1. 如果输入是邮箱/手机号 → 直接用 batch_get_id(最快) - 2. 否则 → 遍历部门查找(需要 contact:department.base:readonly) - 3. 如果部门遍历也失败 → 提示用户改用邮箱/手机号 - """ - print(f" 搜索用户:{name} ...", file=sys.stderr) - - # 方法 1:邮箱/手机号直接查找 - user = _find_user_by_contact(name, config) - if user: - return user - - # 方法 2:部门遍历 - user = _find_user_by_department(name, config) - if user: - return user - - # 都失败 - print(f"\n ❌ 未能找到用户 {name}", file=sys.stderr) - print(f" 建议:", file=sys.stderr) - print(f" 1. 确认已开通 contact:department.base:readonly 权限", file=sys.stderr) - print(f" 2. 改用邮箱搜索:--name user@company.com", file=sys.stderr) - print(f" 3. 改用手机号搜索:--name +8613800138000", file=sys.stderr) - return None - - -# ─── 消息记录 ───────────────────────────────────────────────────────────────── - -def get_chats_with_user(user_open_id: str, config: dict) -> list: - """找到 bot 和目标用户共同在的群聊""" - print(" 获取群聊列表 ...", file=sys.stderr) - - chats = [] - page_token = None - - while True: - params = {"page_size": 100} - if page_token: - params["page_token"] = page_token - - data = api_get("/im/v1/chats", params, config) - if data.get("code") != 0: - print(f" 获取群聊失败:{data.get('msg')}", file=sys.stderr) - break - - items = data.get("data", {}).get("items", []) - chats.extend(items) - - if not data.get("data", {}).get("has_more"): - break - page_token = data.get("data", {}).get("page_token") - - print(f" 共 {len(chats)} 个群聊,检查成员 ...", file=sys.stderr) - - # 过滤:目标用户在其中的群 - result = [] - for chat in chats: - chat_id = chat.get("chat_id") - if not chat_id: - continue - - members_data = api_get( - f"/im/v1/chats/{chat_id}/members", - {"page_size": 100}, - config, - ) - members = members_data.get("data", {}).get("items", []) - for m in members: - if m.get("member_id") == user_open_id or m.get("open_id") == user_open_id: - result.append(chat) - print(f" ✓ {chat.get('name', chat_id)}", file=sys.stderr) - break - - return result - - -def fetch_messages_from_chat( - chat_id: str, - user_open_id: str, - limit: int, - config: dict, -) -> list: - """从指定群聊拉取目标用户的消息""" - messages = [] - page_token = None - - while len(messages) < limit: - params = { - "container_id_type": "chat", - "container_id": chat_id, - "page_size": 50, - "sort_type": "ByCreateTimeDesc", - } - if page_token: - params["page_token"] = page_token - - data = api_get("/im/v1/messages", params, config) - if data.get("code") != 0: - break - - items = data.get("data", {}).get("items", []) - if not items: - break - - for item in items: - sender = item.get("sender", {}) - sender_id = sender.get("id") or sender.get("open_id", "") - if sender_id != user_open_id: - continue - - # 解析消息内容 - content_raw = item.get("body", {}).get("content", "") - try: - content_obj = json.loads(content_raw) - # 富文本消息 - if isinstance(content_obj, dict): - text_parts = [] - for line in content_obj.get("content", []): - for seg in line: - if seg.get("tag") in ("text", "a"): - text_parts.append(seg.get("text", "")) - content = " ".join(text_parts) - else: - content = str(content_obj) - except Exception: - content = content_raw - - content = content.strip() - if not content or content in ("[图片]", "[文件]", "[表情]", "[语音]"): - continue - - ts = item.get("create_time", "") - if ts: - try: - ts = datetime.fromtimestamp(int(ts) / 1000).strftime("%Y-%m-%d %H:%M") - except Exception: - pass - - messages.append({"content": content, "time": ts}) - - if not data.get("data", {}).get("has_more"): - break - page_token = data.get("data", {}).get("page_token") - - return messages[:limit] - - -def fetch_p2p_messages( - chat_id: str, - user_open_id: str, - limit: int, - config: dict, -) -> list: - """使用 user_access_token 从私聊会话拉取消息(包含双方所有消息)""" - messages = [] - page_token = None - - while len(messages) < limit: - params = { - "container_id_type": "chat", - "container_id": chat_id, - "page_size": 50, - "sort_type": "ByCreateTimeDesc", - } - if page_token: - params["page_token"] = page_token - - data = api_get("/im/v1/messages", params, config, use_user_token=True) - if data.get("code") != 0: - print(f" 拉取私聊消息失败(code={data.get('code')}):{data.get('msg')}", file=sys.stderr) - break - - items = data.get("data", {}).get("items", []) - if not items: - break - - for item in items: - sender = item.get("sender", {}) - sender_id = sender.get("id") or sender.get("open_id", "") - - # 解析消息内容 - content_raw = item.get("body", {}).get("content", "") - try: - content_obj = json.loads(content_raw) - if isinstance(content_obj, dict): - # 纯文本消息 - if "text" in content_obj: - content = content_obj["text"] - else: - # 富文本消息 - text_parts = [] - for line in content_obj.get("content", []): - for seg in line: - if seg.get("tag") in ("text", "a"): - text_parts.append(seg.get("text", "")) - content = " ".join(text_parts) - else: - content = str(content_obj) - except Exception: - content = content_raw - - content = content.strip() - if not content or content in ("[图片]", "[文件]", "[表情]", "[语音]"): - continue - - ts = item.get("create_time", "") - if ts: - try: - ts = datetime.fromtimestamp(int(ts) / 1000).strftime("%Y-%m-%d %H:%M") - except Exception: - pass - - is_target = (sender_id == user_open_id) - messages.append({ - "content": content, - "time": ts, - "sender_id": sender_id, - "is_target": is_target, - }) - - if not data.get("data", {}).get("has_more"): - break - page_token = data.get("data", {}).get("page_token") - - return messages[:limit] - - -def collect_messages( - user: dict, - msg_limit: int, - config: dict, -) -> str: - """采集目标用户的所有消息记录(群聊 + 私聊)""" - user_open_id = user.get("open_id") or user.get("user_id", "") - name = user.get("name", "") - - all_messages = [] - chat_sources = [] - - # ── 私聊采集(需要 user_access_token + p2p_chat_id)── - p2p_chat_id = config.get("p2p_chat_id", "") - user_token = config.get("user_access_token", "") - - if user_token and p2p_chat_id: - print(f" 📱 采集私聊消息(chat_id: {p2p_chat_id})...", file=sys.stderr) - p2p_msgs = fetch_p2p_messages(p2p_chat_id, user_open_id, msg_limit, config) - for m in p2p_msgs: - m["chat"] = "私聊" - all_messages.extend(p2p_msgs) - chat_sources.append(f"私聊({len(p2p_msgs)} 条)") - print(f" 获取 {len(p2p_msgs)} 条私聊消息", file=sys.stderr) - elif user_token and not p2p_chat_id: - print(f" ⚠️ 有 user_access_token 但未配置 p2p_chat_id,跳过私聊采集", file=sys.stderr) - print(f" 请在配置中添加 p2p_chat_id(通过发送消息 API 返回值获取)", file=sys.stderr) - - # ── 群聊采集(使用 tenant_access_token)── - remaining = msg_limit - len(all_messages) - if remaining > 0: - chats = get_chats_with_user(user_open_id, config) - if chats: - per_chat_limit = max(100, remaining // len(chats)) - for chat in chats: - chat_id = chat.get("chat_id") - chat_name = chat.get("name", chat_id) - print(f" 拉取「{chat_name}」消息 ...", file=sys.stderr) - - msgs = fetch_messages_from_chat(chat_id, user_open_id, per_chat_limit, config) - for m in msgs: - m["chat"] = chat_name - all_messages.extend(msgs) - chat_sources.append(f"{chat_name}({len(msgs)} 条)") - print(f" 获取 {len(msgs)} 条", file=sys.stderr) - - if not all_messages: - tips = f"# 消息记录\n\n未找到 {name} 的消息记录。\n\n" - tips += "可能原因:\n" - tips += " - 群聊采集:bot 未被添加到相关群聊\n" - tips += " - 私聊采集:未配置 user_access_token 或 p2p_chat_id\n" - tips += "\n私聊采集配置方法:\n" - tips += " 1. 在飞书开放平台开通 im:message 和 im:chat 用户权限\n" - tips += " 2. 通过 OAuth 授权获取 user_access_token(--exchange-code)\n" - tips += " 3. 配置 p2p_chat_id(私聊会话 ID)\n" - return tips - - # 分类输出 - # 私聊消息包含双方对话,标注发言人 - target_msgs = [m for m in all_messages if m.get("is_target", True)] - other_msgs = [m for m in all_messages if not m.get("is_target", True)] - - long_msgs = [m for m in target_msgs if len(m.get("content", "")) > 50] - short_msgs = [m for m in target_msgs if len(m.get("content", "")) <= 50] - - lines = [ - f"# 飞书消息记录(自动采集)", - f"目标:{name}", - f"来源:{', '.join(chat_sources)}", - f"共 {len(all_messages)} 条消息(目标用户 {len(target_msgs)} 条,对话方 {len(other_msgs)} 条)", - "", - "---", - "", - "## 长消息(观点/决策/技术类)", - "", - ] - for m in long_msgs: - lines.append(f"[{m.get('time', '')}][{m.get('chat', '')}] {m['content']}") - lines.append("") - - lines += ["---", "", "## 日常消息(风格参考)", ""] - for m in short_msgs[:300]: - lines.append(f"[{m.get('time', '')}] {m['content']}") - - # 私聊对话上下文(保留双方对话,便于理解语境) - p2p_msgs = [m for m in all_messages if m.get("chat") == "私聊"] - if p2p_msgs: - lines += ["", "---", "", "## 私聊对话上下文(含双方消息)", ""] - # 按时间正序 - p2p_sorted = sorted(p2p_msgs, key=lambda x: x.get("time", "")) - for m in p2p_sorted[:500]: - who = f"[{name}]" if m.get("is_target") else "[对方]" - lines.append(f"[{m.get('time', '')}] {who} {m['content']}") - - return "\n".join(lines) - - -# ─── 文档采集 ───────────────────────────────────────────────────────────────── - -def search_docs_by_user(user_open_id: str, name: str, doc_limit: int, config: dict) -> list: - """搜索目标用户创建或编辑的文档""" - print(f" 搜索 {name} 的文档 ...", file=sys.stderr) - - data = api_post( - "/search/v2/message", - { - "query": name, - "search_type": "docs", - "docs_options": { - "creator_ids": [user_open_id], - }, - "page_size": doc_limit, - }, - config, - ) - - if data.get("code") != 0: - # fallback:用关键词搜索 - print(f" 按创建人搜索失败,改用关键词搜索 ...", file=sys.stderr) - data = api_post( - "/search/v2/message", - { - "query": name, - "search_type": "docs", - "page_size": doc_limit, - }, - config, - ) - - docs = [] - for item in data.get("data", {}).get("results", []): - doc_info = item.get("docs_info", {}) - if doc_info: - docs.append({ - "title": doc_info.get("title", ""), - "url": doc_info.get("url", ""), - "type": doc_info.get("docs_type", ""), - "creator": doc_info.get("creator", {}).get("name", ""), - }) - - print(f" 找到 {len(docs)} 篇文档", file=sys.stderr) - return docs - - -def fetch_doc_content(doc_token: str, doc_type: str, config: dict) -> str: - """拉取单篇文档内容""" - if doc_type in ("doc", "docx"): - data = api_get(f"/docx/v1/documents/{doc_token}/raw_content", {}, config) - return data.get("data", {}).get("content", "") - - elif doc_type == "wiki": - # 先获取 wiki node 信息 - node_data = api_get(f"/wiki/v2/spaces/get_node", {"token": doc_token}, config) - obj_token = node_data.get("data", {}).get("node", {}).get("obj_token", doc_token) - obj_type = node_data.get("data", {}).get("node", {}).get("obj_type", "docx") - return fetch_doc_content(obj_token, obj_type, config) - - return "" - - -def collect_docs(user: dict, doc_limit: int, config: dict) -> str: - """采集目标用户的文档""" - import re - user_open_id = user.get("open_id") or user.get("user_id", "") - name = user.get("name", "") - - docs = search_docs_by_user(user_open_id, name, doc_limit, config) - if not docs: - return f"# 文档内容\n\n未找到 {name} 相关文档\n" - - lines = [ - f"# 文档内容(自动采集)", - f"目标:{name}", - f"共 {len(docs)} 篇", - "", - ] - - for doc in docs: - url = doc.get("url", "") - title = doc.get("title", "无标题") - doc_type = doc.get("type", "") - - print(f" 拉取文档:{title} ...", file=sys.stderr) - - # 从 URL 提取 token - token_match = re.search(r"/(?:wiki|docx|docs|sheets|base)/([A-Za-z0-9]+)", url) - if not token_match: - continue - doc_token = token_match.group(1) - - content = fetch_doc_content(doc_token, doc_type or "docx", config) - if not content or len(content.strip()) < 20: - print(f" 内容为空,跳过", file=sys.stderr) - continue - - lines += [ - f"---", - f"## 《{title}》", - f"链接:{url}", - f"创建人:{doc.get('creator', '')}", - "", - content.strip(), - "", - ] - - return "\n".join(lines) - - -# ─── 多维表格 ───────────────────────────────────────────────────────────────── - -def collect_bitable(app_token: str, config: dict) -> str: - """拉取多维表格内容""" - # 获取所有 table - data = api_get(f"/bitable/v1/apps/{app_token}/tables", {"page_size": 100}, config) - tables = data.get("data", {}).get("items", []) - - if not tables: - return "(多维表格为空)\n" - - lines = [] - for table in tables: - table_id = table.get("table_id") - table_name = table.get("name", table_id) - - # 获取字段 - fields_data = api_get( - f"/bitable/v1/apps/{app_token}/tables/{table_id}/fields", - {"page_size": 100}, - config, - ) - fields = [f.get("field_name", "") for f in fields_data.get("data", {}).get("items", [])] - - # 获取记录 - records_data = api_get( - f"/bitable/v1/apps/{app_token}/tables/{table_id}/records", - {"page_size": 100}, - config, - ) - records = records_data.get("data", {}).get("items", []) - - lines.append(f"### 表:{table_name}") - lines.append("") - lines.append("| " + " | ".join(fields) + " |") - lines.append("| " + " | ".join(["---"] * len(fields)) + " |") - - for rec in records: - row_data = rec.get("fields", {}) - row = [] - for f in fields: - val = row_data.get(f, "") - if isinstance(val, list): - val = " ".join( - v.get("text", str(v)) if isinstance(v, dict) else str(v) - for v in val - ) - row.append(str(val).replace("|", "|").replace("\n", " ")) - lines.append("| " + " | ".join(row) + " |") - - lines.append("") - - return "\n".join(lines) - - -# ─── 主流程 ─────────────────────────────────────────────────────────────────── - -def collect_all( - name: str, - output_dir: Path, - msg_limit: int, - doc_limit: int, - config: dict, -) -> dict: - """采集某同事的所有可用数据,输出到 output_dir""" - output_dir.mkdir(parents=True, exist_ok=True) - results = {} - - print(f"\n🔍 开始采集:{name}\n", file=sys.stderr) - - # Step 1: 搜索用户 - user = find_user(name, config) - if not user: - print(f"❌ 未找到用户 {name},请检查姓名是否正确", file=sys.stderr) - sys.exit(1) - - # Step 2: 采集消息记录 - print(f"\n📨 采集消息记录(上限 {msg_limit} 条)...", file=sys.stderr) - try: - msg_content = collect_messages(user, msg_limit, config) - msg_path = output_dir / "messages.txt" - msg_path.write_text(msg_content, encoding="utf-8") - results["messages"] = str(msg_path) - print(f" ✅ 消息记录 → {msg_path}", file=sys.stderr) - except Exception as e: - print(f" ⚠️ 消息采集失败:{e}", file=sys.stderr) - - # Step 3: 采集文档 - print(f"\n📄 采集文档(上限 {doc_limit} 篇)...", file=sys.stderr) - try: - doc_content = collect_docs(user, doc_limit, config) - doc_path = output_dir / "docs.txt" - doc_path.write_text(doc_content, encoding="utf-8") - results["docs"] = str(doc_path) - print(f" ✅ 文档内容 → {doc_path}", file=sys.stderr) - except Exception as e: - print(f" ⚠️ 文档采集失败:{e}", file=sys.stderr) - - # 写摘要 - summary = { - "name": name, - "user_id": user.get("user_id", ""), - "open_id": user.get("open_id", ""), - "department": user.get("department_path", []), - "collected_at": datetime.now(timezone.utc).isoformat(), - "files": results, - } - (output_dir / "collection_summary.json").write_text( - json.dumps(summary, ensure_ascii=False, indent=2) - ) - - print(f"\n✅ 采集完成,输出目录:{output_dir}", file=sys.stderr) - return results - - -def main() -> None: - parser = argparse.ArgumentParser(description="飞书数据自动采集器") - parser.add_argument("--setup", action="store_true", help="初始化配置") - parser.add_argument("--name", help="同事姓名") - parser.add_argument("--output-dir", default=None, help="输出目录(默认 ./knowledge/{name})") - parser.add_argument("--msg-limit", type=int, default=1000, help="最多采集消息条数(默认 1000)") - parser.add_argument("--doc-limit", type=int, default=20, help="最多采集文档篇数(默认 20)") - parser.add_argument("--exchange-code", metavar="CODE", help="用 OAuth 授权码换取 user_access_token 并保存到配置") - parser.add_argument("--user-token", metavar="TOKEN", help="直接指定 user_access_token(覆盖配置文件)") - parser.add_argument("--p2p-chat-id", metavar="CHAT_ID", help="私聊会话 ID(覆盖配置文件)") - parser.add_argument("--open-id", metavar="OPEN_ID", help="直接指定目标用户的 open_id(跳过用户搜索)") - - args = parser.parse_args() - - if args.setup: - setup_config() - return - - config = load_config() - - # 换取 user_access_token - if args.exchange_code: - token_data = exchange_code_for_token(args.exchange_code, config) - if token_data: - config["user_access_token"] = token_data["access_token"] - config["refresh_token"] = token_data.get("refresh_token", "") - save_config(config) - print(f"✅ user_access_token 已保存(scope: {token_data.get('scope', '')})") - print(f" token: {token_data['access_token'][:20]}...") - else: - print("❌ 换取失败,请检查 code 是否有效") - return - - if not args.name and not args.open_id: - parser.error("请提供 --name 或 --open-id") - - # 命令行参数覆盖配置 - if args.user_token: - config["user_access_token"] = args.user_token - if args.p2p_chat_id: - config["p2p_chat_id"] = args.p2p_chat_id - - output_dir = Path(args.output_dir) if args.output_dir else Path(f"./knowledge/{args.name or 'target'}") - - # 如果提供了 open_id,跳过用户搜索 - if args.open_id: - user = {"open_id": args.open_id, "name": args.name or "target"} - output_dir.mkdir(parents=True, exist_ok=True) - print(f"\n🔍 使用指定 open_id: {args.open_id}\n", file=sys.stderr) - - # 只采集消息 - print(f"📨 采集消息记录(上限 {args.msg_limit} 条)...", file=sys.stderr) - msg_content = collect_messages(user, args.msg_limit, config) - msg_path = output_dir / "messages.txt" - msg_path.write_text(msg_content, encoding="utf-8") - print(f" ✅ 消息记录 → {msg_path}", file=sys.stderr) - return - - collect_all( - name=args.name, - output_dir=output_dir, - msg_limit=args.msg_limit, - doc_limit=args.doc_limit, - config=config, - ) - - -if __name__ == "__main__": - main() diff --git a/tools/feishu_browser.py b/tools/feishu_browser.py deleted file mode 100644 index d46a9d3e..00000000 --- a/tools/feishu_browser.py +++ /dev/null @@ -1,374 +0,0 @@ -#!/usr/bin/env python3 -""" -飞书浏览器抓取器(Playwright 方案) - -复用本机 Chrome 登录态,无需任何 token,能访问你有权限的所有飞书内容。 - -支持: - - 飞书文档(docx/docs) - - 飞书知识库(wiki) - - 飞书表格(sheets)→ 导出为 CSV - - 飞书消息记录(指定群聊) - -安装: - pip install playwright - playwright install chromium - -用法: - python3 feishu_browser.py --url "https://xxx.feishu.cn/wiki/xxx" --output out.txt - python3 feishu_browser.py --url "https://xxx.feishu.cn/docx/xxx" --output out.txt - python3 feishu_browser.py --chat "后端组" --target "张三" --limit 500 --output out.txt - python3 feishu_browser.py --url "https://xxx.feishu.cn/sheets/xxx" --output out.csv -""" - -from __future__ import annotations - -import sys -import time -import json -import argparse -import platform -from pathlib import Path -from typing import Optional - - -def get_default_chrome_profile() -> str: - """根据操作系统返回 Chrome 默认 Profile 路径""" - system = platform.system() - if system == "Darwin": - return str(Path.home() / "Library/Application Support/Google/Chrome/Default") - elif system == "Linux": - return str(Path.home() / ".config/google-chrome/Default") - elif system == "Windows": - import os - return str(Path(os.environ.get("LOCALAPPDATA", "")) / "Google/Chrome/User Data/Default") - return str(Path.home() / ".config/google-chrome/Default") - - -def make_context(playwright, chrome_profile: Optional[str], headless: bool): - """创建复用登录态的浏览器上下文""" - profile = chrome_profile or get_default_chrome_profile() - try: - ctx = playwright.chromium.launch_persistent_context( - user_data_dir=profile, - headless=headless, - args=[ - "--disable-blink-features=AutomationControlled", - "--no-first-run", - "--no-default-browser-check", - ], - ignore_default_args=["--enable-automation"], - viewport={"width": 1280, "height": 900}, - ) - return ctx - except Exception as e: - print(f"⚠️ 无法加载 Chrome Profile:{e}", file=sys.stderr) - print(f" 尝试的路径:{profile}", file=sys.stderr) - print(" 请用 --chrome-profile 手动指定路径", file=sys.stderr) - sys.exit(1) - - -def detect_page_type(url: str) -> str: - """根据 URL 判断飞书页面类型""" - if "/wiki/" in url: - return "wiki" - elif "/docx/" in url or "/docs/" in url: - return "doc" - elif "/sheets/" in url or "/spreadsheets/" in url: - return "sheet" - elif "/base/" in url: - return "base" - else: - return "unknown" - - -def fetch_doc(page, url: str) -> str: - """抓取飞书文档或 Wiki 的文本内容""" - page.goto(url, wait_until="domcontentloaded", timeout=30000) - - # 等待编辑器加载(飞书文档渲染较慢) - selectors = [ - ".docs-reader-content", - ".lark-editor-content", - "[data-block-type]", - ".doc-render-core", - ".wiki-content", - ".node-doc-content", - ] - - loaded = False - for sel in selectors: - try: - page.wait_for_selector(sel, timeout=15000) - loaded = True - break - except Exception: - continue - - if not loaded: - # 等待一段时间后直接提取 body 文本 - time.sleep(5) - - # 额外等待异步内容渲染 - time.sleep(2) - - # 尝试多个选择器提取正文 - for sel in selectors: - try: - el = page.query_selector(sel) - if el: - text = el.inner_text() - if len(text.strip()) > 50: - return text.strip() - except Exception: - continue - - # fallback:提取整个 body - text = page.inner_text("body") - return text.strip() - - -def fetch_sheet(page, url: str) -> str: - """抓取飞书表格,转为 CSV 格式""" - page.goto(url, wait_until="domcontentloaded", timeout=30000) - - try: - page.wait_for_selector(".spreadsheet-container, .sheet-container", timeout=15000) - except Exception: - time.sleep(5) - - time.sleep(3) - - # 通过 JS 提取表格数据 - data = page.evaluate(""" - () => { - const rows = []; - // 尝试从 DOM 提取可见单元格 - const cells = document.querySelectorAll('[data-row][data-col]'); - if (cells.length === 0) return null; - - const grid = {}; - let maxRow = 0, maxCol = 0; - cells.forEach(cell => { - const r = parseInt(cell.getAttribute('data-row')); - const c = parseInt(cell.getAttribute('data-col')); - if (!grid[r]) grid[r] = {}; - grid[r][c] = cell.innerText.replace(/\\n/g, ' ').trim(); - maxRow = Math.max(maxRow, r); - maxCol = Math.max(maxCol, c); - }); - - for (let r = 0; r <= maxRow; r++) { - const row = []; - for (let c = 0; c <= maxCol; c++) { - row.push(grid[r] && grid[r][c] ? grid[r][c] : ''); - } - rows.push(row); - } - return rows; - } - """) - - if data: - lines = [] - for row in data: - lines.append(",".join(f'"{cell}"' for cell in row)) - return "\n".join(lines) - - # fallback:直接提取文本 - return page.inner_text("body") - - -def fetch_messages(page, chat_name: str, target_name: str, limit: int = 500) -> str: - """ - 抓取指定群聊中目标人物的消息记录。 - 需要先导航到飞书 Web 版消息页面。 - """ - # 打开飞书消息页 - page.goto("https://applink.feishu.cn/client/chat/open", wait_until="domcontentloaded", timeout=20000) - time.sleep(3) - - # 尝试搜索群聊 - try: - # 点击搜索 - search_btn = page.query_selector('[data-test-id="search-btn"], .search-button, [placeholder*="搜索"]') - if search_btn: - search_btn.click() - time.sleep(1) - page.keyboard.type(chat_name) - time.sleep(2) - - # 选择第一个结果 - result = page.query_selector('.search-result-item:first-child, .im-search-item:first-child') - if result: - result.click() - time.sleep(2) - except Exception as e: - print(f"⚠️ 自动搜索群聊失败:{e}", file=sys.stderr) - print(f" 请手动导航到「{chat_name}」群聊,然后按回车继续...", file=sys.stderr) - input() - - # 向上滚动加载历史消息 - print(f"正在加载消息历史...", file=sys.stderr) - messages_container = page.query_selector('.message-list, .im-message-list, [data-testid="message-list"]') - - if messages_container: - for _ in range(10): # 滚动 10 次 - page.evaluate("el => el.scrollTop = 0", messages_container) - time.sleep(1.5) - else: - for _ in range(10): - page.keyboard.press("Control+Home") - time.sleep(1.5) - - time.sleep(2) - - # 提取消息 - messages = page.evaluate(f""" - () => {{ - const target = "{target_name}"; - const results = []; - - // 常见的消息 DOM 结构 - const msgSelectors = [ - '.message-item', - '.im-message-item', - '[data-message-id]', - '.msg-list-item', - ]; - - let items = []; - for (const sel of msgSelectors) {{ - items = document.querySelectorAll(sel); - if (items.length > 0) break; - }} - - items.forEach(item => {{ - const senderEl = item.querySelector( - '.sender-name, .message-sender, [data-testid="sender-name"], .name' - ); - const contentEl = item.querySelector( - '.message-content, .msg-content, [data-testid="message-content"], .text-content' - ); - const timeEl = item.querySelector( - '.message-time, .msg-time, [data-testid="message-time"], .time' - ); - - const sender = senderEl ? senderEl.innerText.trim() : ''; - const content = contentEl ? contentEl.innerText.trim() : ''; - const time = timeEl ? timeEl.innerText.trim() : ''; - - if (!content) return; - if (target && !sender.includes(target)) return; - - results.push({{ sender, content, time }}); - }}); - - return results.slice(-{limit}); - }} - """) - - if not messages: - print("⚠️ 未能自动提取消息,尝试提取页面文本", file=sys.stderr) - return page.inner_text("body") - - # 按权重分类输出 - long_msgs = [m for m in messages if len(m.get("content", "")) > 50] - short_msgs = [m for m in messages if len(m.get("content", "")) <= 50] - - lines = [ - f"# 飞书消息记录(浏览器抓取)", - f"群聊:{chat_name}", - f"目标人物:{target_name}", - f"共 {len(messages)} 条消息", - "", - "---", - "", - "## 长消息(观点/决策类)", - "", - ] - for m in long_msgs: - lines.append(f"[{m.get('time', '')}] {m.get('content', '')}") - lines.append("") - - lines += ["---", "", "## 日常消息", ""] - for m in short_msgs[:200]: - lines.append(f"[{m.get('time', '')}] {m.get('content', '')}") - - return "\n".join(lines) - - -def main() -> None: - parser = argparse.ArgumentParser(description="飞书浏览器抓取器(复用 Chrome 登录态)") - parser.add_argument("--url", help="飞书文档/Wiki/表格链接") - parser.add_argument("--chat", help="群聊名称(抓取消息记录时使用)") - parser.add_argument("--target", help="目标人物姓名(只提取此人的消息)") - parser.add_argument("--limit", type=int, default=500, help="最多抓取消息条数(默认 500)") - parser.add_argument("--output", default=None, help="输出文件路径(默认打印到 stdout)") - parser.add_argument("--chrome-profile", default=None, help="Chrome Profile 路径(默认自动检测)") - parser.add_argument("--headless", action="store_true", help="无头模式(不显示浏览器窗口)") - parser.add_argument("--show-browser", action="store_true", help="显示浏览器窗口(调试用)") - - args = parser.parse_args() - - if not args.url and not args.chat: - parser.error("请提供 --url(文档链接)或 --chat(群聊名称)") - - try: - from playwright.sync_api import sync_playwright - except ImportError: - print("错误:请先安装 Playwright:pip install playwright && playwright install chromium", file=sys.stderr) - sys.exit(1) - - headless = args.headless and not args.show_browser - - print(f"启动浏览器({'无头' if headless else '有界面'}模式)...", file=sys.stderr) - - with sync_playwright() as p: - ctx = make_context(p, args.chrome_profile, headless=headless) - page = ctx.new_page() - - # 检查是否已登录 - page.goto("https://www.feishu.cn", wait_until="domcontentloaded", timeout=15000) - time.sleep(2) - if "login" in page.url.lower() or "signin" in page.url.lower(): - print("⚠️ 检测到未登录状态。", file=sys.stderr) - print(" 请在打开的浏览器窗口中登录飞书,登录后按回车继续...", file=sys.stderr) - if headless: - print(" 提示:请用 --show-browser 参数显示浏览器窗口以完成登录", file=sys.stderr) - sys.exit(1) - input() - - # 根据任务类型执行 - if args.url: - page_type = detect_page_type(args.url) - print(f"页面类型:{page_type},开始抓取...", file=sys.stderr) - - if page_type == "sheet": - content = fetch_sheet(page, args.url) - else: - content = fetch_doc(page, args.url) - - elif args.chat: - content = fetch_messages( - page, - chat_name=args.chat, - target_name=args.target or "", - limit=args.limit, - ) - - ctx.close() - - if not content or len(content.strip()) < 10: - print("⚠️ 未能提取到有效内容", file=sys.stderr) - sys.exit(1) - - if args.output: - Path(args.output).write_text(content, encoding="utf-8") - print(f"✅ 已保存到 {args.output}({len(content)} 字符)", file=sys.stderr) - else: - print(content) - - -if __name__ == "__main__": - main() diff --git a/tools/feishu_mcp_client.py b/tools/feishu_mcp_client.py deleted file mode 100644 index f9ca2069..00000000 --- a/tools/feishu_mcp_client.py +++ /dev/null @@ -1,314 +0,0 @@ -#!/usr/bin/env python3 -""" -飞书 MCP 客户端封装(cso1z/Feishu-MCP 方案) - -通过 Feishu MCP Server 读取文档、wiki、消息记录。 -适合:公司已授权的文档、有 App token 权限的内容。 - -前置要求: - 1. 安装 Feishu MCP:npm install -g feishu-mcp - 2. 配置 App ID 和 App Secret(飞书开放平台创建企业自建应用) - 3. 给应用开通必要权限(见下方 REQUIRED_PERMISSIONS) - -权限列表(飞书开放平台 → 权限管理 → 开通): - - docs:doc:readonly 读取文档 - - wiki:wiki:readonly 读取知识库 - - im:message:readonly 读取消息 - - bitable:app:readonly 读取多维表格 - - sheets:spreadsheet:readonly 读取表格 - -用法: - # 配置 token(一次性) - python3 feishu_mcp_client.py --setup - - # 读取文档 - python3 feishu_mcp_client.py --url "https://xxx.feishu.cn/wiki/xxx" --output out.txt - - # 读取消息记录 - python3 feishu_mcp_client.py --chat-id "oc_xxx" --target "张三" --output out.txt - - # 列出某空间下的所有文档 - python3 feishu_mcp_client.py --list-wiki --space-id "xxx" -""" - -from __future__ import annotations - -import os -import sys -import json -import argparse -import subprocess -from pathlib import Path -from typing import Optional - - -CONFIG_PATH = Path.home() / ".distilly" / "feishu_config.json" -LEGACY_CONFIG_PATH = Path.home() / ".colleague-skill" / "feishu_config.json" - - -# ─── 配置管理 ──────────────────────────────────────────────────────────────── - -def load_config() -> dict: - if CONFIG_PATH.exists(): - return json.loads(CONFIG_PATH.read_text()) - if LEGACY_CONFIG_PATH.exists(): - return json.loads(LEGACY_CONFIG_PATH.read_text()) - return {} - - -def save_config(config: dict) -> None: - CONFIG_PATH.parent.mkdir(parents=True, exist_ok=True) - CONFIG_PATH.write_text(json.dumps(config, indent=2)) - CONFIG_PATH.chmod(0o600) - print(f"配置已保存到 {CONFIG_PATH}") - - -def setup_config() -> None: - print("=== 飞书 MCP 配置 ===") - print("请前往飞书开放平台(open.feishu.cn)创建企业自建应用,获取以下信息:\n") - - app_id = input("App ID (cli_xxx): ").strip() - app_secret = input("App Secret: ").strip() - - print("\n配置方式选择:") - print(" [1] App Token(应用权限,需要在飞书后台开通对应权限)") - print(" [2] User Token(个人权限,能访问你本人有权限的所有内容,需要定期刷新)") - mode = input("选择 [1/2],默认 1:").strip() or "1" - - config = { - "app_id": app_id, - "app_secret": app_secret, - "mode": "app" if mode == "1" else "user", - } - - if mode == "2": - print("\n获取 User Token:飞书开放平台 → OAuth 2.0 → 获取 user_access_token") - user_token = input("User Access Token (u-xxx):").strip() - config["user_token"] = user_token - print("注意:User Token 有效期约 2 小时,过期后需要重新配置") - - save_config(config) - print("\n✅ 配置完成!") - - -# ─── MCP 调用封装 ───────────────────────────────────────────────────────────── - -def call_mcp(tool: str, params: dict, config: dict) -> dict: - """ - 通过 npx 调用 feishu-mcp 工具。 - feishu-mcp 支持 stdio 模式,直接 JSON 通信。 - """ - env = os.environ.copy() - env["FEISHU_APP_ID"] = config.get("app_id", "") - env["FEISHU_APP_SECRET"] = config.get("app_secret", "") - - if config.get("mode") == "user" and config.get("user_token"): - env["FEISHU_USER_ACCESS_TOKEN"] = config["user_token"] - - payload = json.dumps({ - "jsonrpc": "2.0", - "method": "tools/call", - "params": { - "name": tool, - "arguments": params, - }, - "id": 1, - }) - - try: - result = subprocess.run( - ["npx", "-y", "feishu-mcp", "--stdio"], - input=payload, - capture_output=True, - text=True, - env=env, - timeout=30, - ) - if result.returncode != 0: - raise RuntimeError(f"MCP 调用失败:{result.stderr}") - return json.loads(result.stdout) - except FileNotFoundError: - print("错误:未找到 npx,请先安装 Node.js", file=sys.stderr) - print("安装 Feishu MCP:npm install -g feishu-mcp", file=sys.stderr) - sys.exit(1) - - -def extract_doc_token(url: str) -> tuple[str, str]: - """从飞书 URL 中提取文档 token 和类型""" - import re - patterns = [ - (r"/wiki/([A-Za-z0-9]+)", "wiki"), - (r"/docx/([A-Za-z0-9]+)", "docx"), - (r"/docs/([A-Za-z0-9]+)", "doc"), - (r"/sheets/([A-Za-z0-9]+)", "sheet"), - (r"/base/([A-Za-z0-9]+)", "base"), - ] - for pattern, doc_type in patterns: - m = re.search(pattern, url) - if m: - return m.group(1), doc_type - raise ValueError(f"无法从 URL 解析文档 token:{url}") - - -# ─── 功能函数 ───────────────────────────────────────────────────────────────── - -def fetch_doc_via_mcp(url: str, config: dict) -> str: - """通过 MCP 读取飞书文档或 Wiki""" - token, doc_type = extract_doc_token(url) - - if doc_type == "wiki": - result = call_mcp("get_wiki_node", {"token": token}, config) - elif doc_type in ("docx", "doc"): - result = call_mcp("get_doc_content", {"doc_token": token}, config) - elif doc_type == "sheet": - result = call_mcp("get_spreadsheet_content", {"spreadsheet_token": token}, config) - else: - raise ValueError(f"不支持的文档类型:{doc_type}") - - # 提取 MCP 返回的内容 - if "result" in result: - content = result["result"] - if isinstance(content, list): - # MCP tool result 格式 - for item in content: - if isinstance(item, dict) and item.get("type") == "text": - return item.get("text", "") - elif isinstance(content, str): - return content - elif "error" in result: - raise RuntimeError(f"MCP 返回错误:{result['error']}") - - return json.dumps(result, ensure_ascii=False, indent=2) - - -def fetch_messages_via_mcp( - chat_id: str, - target_name: str, - limit: int, - config: dict, -) -> str: - """通过 MCP 读取群聊消息记录""" - result = call_mcp( - "get_chat_messages", - { - "chat_id": chat_id, - "page_size": min(limit, 50), # 飞书 API 单次最多 50 条 - }, - config, - ) - - messages = [] - raw = result.get("result", []) - if isinstance(raw, list): - messages = raw - elif isinstance(raw, str): - try: - messages = json.loads(raw) - except Exception: - return raw - - # 过滤目标人物 - if target_name: - messages = [ - m for m in messages - if target_name in str(m.get("sender", {}).get("name", "")) - ] - - # 分类输出 - long_msgs = [m for m in messages if len(str(m.get("content", ""))) > 50] - short_msgs = [m for m in messages if len(str(m.get("content", ""))) <= 50] - - lines = [ - "# 飞书消息记录(MCP 方案)", - f"群聊 ID:{chat_id}", - f"目标人物:{target_name or '全部'}", - f"共 {len(messages)} 条", - "", - "---", - "", - "## 长消息", - "", - ] - for m in long_msgs: - sender = m.get("sender", {}).get("name", "") - content = m.get("content", "") - ts = m.get("create_time", "") - lines.append(f"[{ts}] {sender}:{content}") - lines.append("") - - lines += ["---", "", "## 日常消息", ""] - for m in short_msgs[:200]: - sender = m.get("sender", {}).get("name", "") - content = m.get("content", "") - lines.append(f"{sender}:{content}") - - return "\n".join(lines) - - -def list_wiki_docs(space_id: str, config: dict) -> str: - """列出知识库空间下的所有文档""" - result = call_mcp("list_wiki_nodes", {"space_id": space_id}, config) - raw = result.get("result", "") - if isinstance(raw, str): - return raw - return json.dumps(raw, ensure_ascii=False, indent=2) - - -# ─── CLI ───────────────────────────────────────────────────────────────────── - -def main() -> None: - parser = argparse.ArgumentParser(description="飞书 MCP 客户端") - parser.add_argument("--setup", action="store_true", help="初始化配置(App ID / Secret)") - parser.add_argument("--url", help="飞书文档/Wiki/表格链接") - parser.add_argument("--chat-id", help="群聊 ID(oc_xxx 格式)") - parser.add_argument("--target", help="目标人物姓名") - parser.add_argument("--limit", type=int, default=500, help="最多获取消息数") - parser.add_argument("--list-wiki", action="store_true", help="列出知识库文档") - parser.add_argument("--space-id", help="知识库 Space ID") - parser.add_argument("--output", default=None, help="输出文件路径") - - args = parser.parse_args() - - if args.setup: - setup_config() - return - - config = load_config() - if not config: - print("错误:尚未配置,请先运行:python3 feishu_mcp_client.py --setup", file=sys.stderr) - sys.exit(1) - - content = "" - - if args.url: - print(f"通过 MCP 读取:{args.url}", file=sys.stderr) - content = fetch_doc_via_mcp(args.url, config) - - elif args.chat_id: - print(f"通过 MCP 读取消息:{args.chat_id}", file=sys.stderr) - content = fetch_messages_via_mcp( - args.chat_id, - args.target or "", - args.limit, - config, - ) - - elif args.list_wiki: - if not args.space_id: - print("错误:--list-wiki 需要 --space-id", file=sys.stderr) - sys.exit(1) - content = list_wiki_docs(args.space_id, config) - - else: - parser.print_help() - return - - if args.output: - Path(args.output).write_text(content, encoding="utf-8") - print(f"✅ 已保存到 {args.output}", file=sys.stderr) - else: - print(content) - - -if __name__ == "__main__": - main() diff --git a/tools/feishu_parser.py b/tools/feishu_parser.py deleted file mode 100644 index 4c1c54a5..00000000 --- a/tools/feishu_parser.py +++ /dev/null @@ -1,251 +0,0 @@ -#!/usr/bin/env python3 -""" -飞书消息导出 JSON 解析器 - -支持的导出格式: -1. 飞书官方导出(群聊记录):通常为 JSON 数组,每条消息包含 sender、content、timestamp -2. 手动整理的 TXT 格式(每行:时间 发送人:内容) - -用法: - python feishu_parser.py --file messages.json --target "张三" --output output.txt - python feishu_parser.py --file messages.txt --target "张三" --output output.txt -""" - -import json -import re -import sys -import argparse -from pathlib import Path -from datetime import datetime - - -def parse_feishu_json(file_path: str, target_name: str) -> list[dict]: - """解析飞书官方导出的 JSON 格式消息""" - with open(file_path, "r", encoding="utf-8") as f: - data = json.load(f) - - messages = [] - - # 兼容多种 JSON 结构 - if isinstance(data, list): - raw_messages = data - elif isinstance(data, dict): - # 可能在 data.messages 或 data.records 等字段下 - raw_messages = ( - data.get("messages") - or data.get("records") - or data.get("data") - or [] - ) - else: - return [] - - for msg in raw_messages: - sender = ( - msg.get("sender_name") - or msg.get("sender") - or msg.get("from") - or msg.get("user_name") - or "" - ) - content = ( - msg.get("content") - or msg.get("text") - or msg.get("message") - or msg.get("body") - or "" - ) - timestamp = ( - msg.get("timestamp") - or msg.get("create_time") - or msg.get("time") - or "" - ) - - # content 可能是嵌套结构 - if isinstance(content, dict): - content = content.get("text") or content.get("content") or str(content) - if isinstance(content, list): - content = " ".join( - c.get("text", "") if isinstance(c, dict) else str(c) - for c in content - ) - - # 过滤:只保留目标人发送的消息 - if target_name and target_name not in str(sender): - continue - - # 过滤:跳过系统消息、表情包、撤回消息 - if not content or content.strip() in ["[图片]", "[文件]", "[撤回了一条消息]", "[语音]"]: - continue - - messages.append({ - "sender": str(sender), - "content": str(content).strip(), - "timestamp": str(timestamp), - }) - - return messages - - -def parse_feishu_txt(file_path: str, target_name: str) -> list[dict]: - """解析手动整理的 TXT 格式消息(格式:时间 发送人:内容)""" - messages = [] - - with open(file_path, "r", encoding="utf-8") as f: - lines = f.readlines() - - # 匹配格式:2024-01-01 10:00 张三:消息内容 - pattern = re.compile( - r"^(?P