Skip to content

fix: make the map readable for a Chinese project - #67

Merged
TheCrazyAnt merged 3 commits into
mainfrom
fix/chinese-capability-matching
Sep 11, 2026
Merged

TheCrazyAnt merged 3 commits into
mainfrom
fix/chinese-capability-matching

Conversation

@TheCrazyAnt

Copy link
Copy Markdown
Owner

From a real deployment, not a bug report: a Chinese project's feature list read as (底层)创作库事务·导出命名 and (底层)接口登录校验·(底层)登录校验, while the five business capabilities its README declares appeared nowhere. The user's words were 「有点乱,没那么清晰」.

It turned out not to be a layout problem. The map was drawing the wrong things, for four separate reasons — three of which only ever affected Chinese.

1. A two-character Chinese heading was not a capability

isCapabilityLabel required 3 characters — correct for rejecting AI and v2, wrong for Chinese, where a term is complete at two. Measured on a five-heading README:

before:  OK "出点子"  OK "写故事"   MISS "选题"  MISS "成片"  MISS "导出"
after:   all five extracted

2. Han runs were matched greedily

[㐀-鿿]{2,8} swallowed a whole run, so 给选题出点子 became one token and the 8-character cap cut 保险短视频编导工作台 into 保险短视频编导工 + 作台. Neither meets the code's vocabulary.

3. The two sides tokenized differently

The reader's keywords() kept two-character terms; the compiler's semanticTokens() deleted them again with length >= 3. That filter is a no-op for Latin ([a-z][a-z0-9-]{2,} is already ≥3 chars), so it only ever removed Chinese. A term could survive extraction and still fail to match.

Fix for 1–3: one shared splitter (cjkTokens in analysis-kit) keeps a short run whole and emits every 2-gram, so 选题 is findable inside 热点选题. Both sides use it; the Latin floor applies only to Latin.

4. Every internal entry was a feature (not language-specific)

features.ts makes every entrypoint unreachable from a user action its own feature, and sorts by health then alphabetically — not importance. A 744-file project contributes one per internal transaction, auth check and helper, and those sort to the top.

Documented capabilities now lead the list; the rest move behind an "Other entry points" disclosure stating what it holds and how many. Nothing is removed. A project that documents nothing keeps its flat list — splitting on an absent signal would empty the list and hide the whole map, strictly worse than before.

Two problems found by probing my own fix

Neither was reported; both came from testing the change rather than trusting it.

False matches. 2-grams put far more words in reach, and an entry hit weighs 8 — enough to decide a match alone. An entry called 自动重试 matched the capability 成片, because that capability's description contains 自动. Presenting a retry loop as a business capability is worse than a code name: it is wrong and it looks right. Only a capability's own name carries entry weight now.

Capabilities without a route. A project whose logic is exported functions has no entrypoint node at all, so 选题 was rejected by the weak-match floor despite a step called 热点选题. A capability's name inside a step is now strong evidence (weight 3); description words stay at 1 and still cannot clear the floor alone.

Evidence

Red before green for all 7 new tests, with messages matching the diagnosis exactly (expected [ '出点子' ] to deeply equal ArrayContaining ["选题", "成片", "出点子"], expected '生成成片' to be '成片', expected '成片' to be '自动重试').

Mutation proof — each wrong implementation had to turn a test red; sources restored byte-identical afterwards, verified with diff:

Mutation Result
A — re-apply the Latin length floor to Chinese 1 red
B — cjkTokens emits no 2-grams 2 red
C — re-apply filter(length >= 3) in the matcher 1 red

End to end, on a Chinese fixture built with the real CLI:

before:  POST /api/ideas   POST /api/story   POST /api/topics
after:   [文档] 出点子   [文档] 成片   [文档] 选题

Verified in the running viewer with a graph shaped like the reported one (3 documented + 5 internal): the three capabilities lead the sidebar and the five internal entries sit behind the disclosure.

npm run check: typecheck, 24 files / 221 tests (213 before, 8 new), build — all pass. Also added the missing analysis-kit project reference to packages/project-reader/tsconfig.json, without which the new import passes vitest but fails the build.

🤖 Generated with Claude Code

TheCrazyAnt and others added 3 commits September 11, 2026 13:32
A Chinese project's README could declare its capabilities and have none of them
reach the map. Three places applied rules written for Latin text to Han
characters, and each one alone was enough to break the match.

The capability filter required three characters, a floor that exists to reject
"AI" and "v2". A Chinese term is complete at two, so 选题, 成片, 导出 and 登录
were dropped before anything could match them — in a five-heading fixture only
the three-character ones survived.

Han runs were then matched greedily, which produced tokens no reader would
search for: a whole clause ("给选题出点子"), a function word welded to its term
("从热点选题"), and an eight-character cap slicing 工作台 into 工 + 作台.

Worst of the three, the two sides disagreed. The reader kept two-character
terms; the matcher filtered them out again, so even a term that survived
extraction could not meet the code's own vocabulary.

Both sides now share one splitter — the whole run when it is short enough to be
a term, plus every 2-gram — and the Latin floor applies only to Latin. The
meaningless grams this admits are the deliberate trade: a spurious token can
only fail to match, while a missing one loses the capability outright.

Proven by mutation: restoring the Latin floor, dropping the 2-grams, or
re-applying the length filter each turns the new tests red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Splitting Han runs into 2-grams put far more words within reach of the matcher,
including words that come from a capability's description rather than its name.
An entry-term hit weighs 8 and so decides the match on its own, which made a
new failure possible: an entry called 自动重试 matched the capability 成片,
because that capability's description reads 自动拍成可导进剪映的成片.

Presenting a retry loop as a business capability is worse than leaving it named
after code — it is wrong, and it looks right, which is the one failure mode this
tool exists to avoid.

The capability's name alone now carries entry weight. Its description still
contributes step evidence, where it is weighed at 1 and has to clear the
existing weak-match floor.

Found by probing the fix that preceded it, not by a report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every entrypoint a user action cannot reach becomes its own feature, so a large
codebase contributes one per internal transaction, auth check and helper. They
are real entries, but the list sorts by health and then alphabetically — not by
importance — so in a 744-file project the internal ones surfaced first and the
reader scrolled past the capabilities the product actually offers.

`product` already carries the distinction: it is set only when a documented
capability matched strongly enough to lend its name, which is exactly "the
project says this exists". Documented capabilities now lead the list; the rest
move behind a disclosure that says what it holds and how many. Nothing is
removed and nothing is hidden.

A project that documents nothing would otherwise get an empty list and a drawer
containing its entire map, which is strictly worse than the flat list it had.
With no signal to separate on, the flat list is kept.

Matching also had to reach capabilities that arrive without a route: a project
whose business logic is exported functions has no entrypoint node to carry the
high-weight entry match, so 选题 was rejected even with a step called 热点选题.
A capability's own name inside a step is now strong evidence, weighed above the
words of its description, which stay weak enough that they still cannot clear
the weak-match floor alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@TheCrazyAnt
TheCrazyAnt merged commit 34c9c7d into main Sep 11, 2026
8 checks passed
@TheCrazyAnt
TheCrazyAnt deleted the fix/chinese-capability-matching branch September 11, 2026 07:47
@TheCrazyAnt TheCrazyAnt mentioned this pull request Sep 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant