fix: make the map readable for a Chinese project - #67
Merged
Merged
Conversation
A Chinese project's README could declare its capabilities and have none of them
reach the map. Three places applied rules written for Latin text to Han
characters, and each one alone was enough to break the match.
The capability filter required three characters, a floor that exists to reject
"AI" and "v2". A Chinese term is complete at two, so 选题, 成片, 导出 and 登录
were dropped before anything could match them — in a five-heading fixture only
the three-character ones survived.
Han runs were then matched greedily, which produced tokens no reader would
search for: a whole clause ("给选题出点子"), a function word welded to its term
("从热点选题"), and an eight-character cap slicing 工作台 into 工 + 作台.
Worst of the three, the two sides disagreed. The reader kept two-character
terms; the matcher filtered them out again, so even a term that survived
extraction could not meet the code's own vocabulary.
Both sides now share one splitter — the whole run when it is short enough to be
a term, plus every 2-gram — and the Latin floor applies only to Latin. The
meaningless grams this admits are the deliberate trade: a spurious token can
only fail to match, while a missing one loses the capability outright.
Proven by mutation: restoring the Latin floor, dropping the 2-grams, or
re-applying the length filter each turns the new tests red.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Splitting Han runs into 2-grams put far more words within reach of the matcher, including words that come from a capability's description rather than its name. An entry-term hit weighs 8 and so decides the match on its own, which made a new failure possible: an entry called 自动重试 matched the capability 成片, because that capability's description reads 自动拍成可导进剪映的成片. Presenting a retry loop as a business capability is worse than leaving it named after code — it is wrong, and it looks right, which is the one failure mode this tool exists to avoid. The capability's name alone now carries entry weight. Its description still contributes step evidence, where it is weighed at 1 and has to clear the existing weak-match floor. Found by probing the fix that preceded it, not by a report. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every entrypoint a user action cannot reach becomes its own feature, so a large codebase contributes one per internal transaction, auth check and helper. They are real entries, but the list sorts by health and then alphabetically — not by importance — so in a 744-file project the internal ones surfaced first and the reader scrolled past the capabilities the product actually offers. `product` already carries the distinction: it is set only when a documented capability matched strongly enough to lend its name, which is exactly "the project says this exists". Documented capabilities now lead the list; the rest move behind a disclosure that says what it holds and how many. Nothing is removed and nothing is hidden. A project that documents nothing would otherwise get an empty list and a drawer containing its entire map, which is strictly worse than the flat list it had. With no signal to separate on, the flat list is kept. Matching also had to reach capabilities that arrive without a route: a project whose business logic is exported functions has no entrypoint node to carry the high-weight entry match, so 选题 was rejected even with a step called 热点选题. A capability's own name inside a step is now strong evidence, weighed above the words of its description, which stay weak enough that they still cannot clear the weak-match floor alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
From a real deployment, not a bug report: a Chinese project's feature list read as
(底层)创作库事务·导出命名and(底层)接口登录校验·(底层)登录校验, while the five business capabilities its README declares appeared nowhere. The user's words were 「有点乱,没那么清晰」.It turned out not to be a layout problem. The map was drawing the wrong things, for four separate reasons — three of which only ever affected Chinese.
1. A two-character Chinese heading was not a capability
isCapabilityLabelrequired 3 characters — correct for rejectingAIandv2, wrong for Chinese, where a term is complete at two. Measured on a five-heading README:2. Han runs were matched greedily
[㐀-鿿]{2,8}swallowed a whole run, so给选题出点子became one token and the 8-character cap cut保险短视频编导工作台into保险短视频编导工+作台. Neither meets the code's vocabulary.3. The two sides tokenized differently
The reader's
keywords()kept two-character terms; the compiler'ssemanticTokens()deleted them again withlength >= 3. That filter is a no-op for Latin ([a-z][a-z0-9-]{2,}is already ≥3 chars), so it only ever removed Chinese. A term could survive extraction and still fail to match.Fix for 1–3: one shared splitter (
cjkTokensinanalysis-kit) keeps a short run whole and emits every 2-gram, so选题is findable inside热点选题. Both sides use it; the Latin floor applies only to Latin.4. Every internal entry was a feature (not language-specific)
features.tsmakes every entrypoint unreachable from a user action its own feature, and sorts by health then alphabetically — not importance. A 744-file project contributes one per internal transaction, auth check and helper, and those sort to the top.Documented capabilities now lead the list; the rest move behind an "Other entry points" disclosure stating what it holds and how many. Nothing is removed. A project that documents nothing keeps its flat list — splitting on an absent signal would empty the list and hide the whole map, strictly worse than before.
Two problems found by probing my own fix
Neither was reported; both came from testing the change rather than trusting it.
False matches. 2-grams put far more words in reach, and an entry hit weighs 8 — enough to decide a match alone. An entry called
自动重试matched the capability成片, because that capability's description contains自动. Presenting a retry loop as a business capability is worse than a code name: it is wrong and it looks right. Only a capability's own name carries entry weight now.Capabilities without a route. A project whose logic is exported functions has no entrypoint node at all, so
选题was rejected by the weak-match floor despite a step called热点选题. A capability's name inside a step is now strong evidence (weight 3); description words stay at 1 and still cannot clear the floor alone.Evidence
Red before green for all 7 new tests, with messages matching the diagnosis exactly (
expected [ '出点子' ] to deeply equal ArrayContaining ["选题", "成片", "出点子"],expected '生成成片' to be '成片',expected '成片' to be '自动重试').Mutation proof — each wrong implementation had to turn a test red; sources restored byte-identical afterwards, verified with
diff:cjkTokensemits no 2-gramsfilter(length >= 3)in the matcherEnd to end, on a Chinese fixture built with the real CLI:
Verified in the running viewer with a graph shaped like the reported one (3 documented + 5 internal): the three capabilities lead the sidebar and the five internal entries sit behind the disclosure.
npm run check: typecheck, 24 files / 221 tests (213 before, 8 new), build — all pass. Also added the missinganalysis-kitproject reference topackages/project-reader/tsconfig.json, without which the new import passes vitest but fails the build.🤖 Generated with Claude Code