Repository navigation
Claude/fiction analysis tool vb1p h - #56
Conversation
…plied blindly
Weak verbs: the detector now measures this manuscript's reliance on each
generic verb against the family of specific verbs the author already uses,
and stays silent when the author varies. Where it speaks it quotes the count
and rate. Participles (", made quietly"), passives, aspectual uses
("started to"), clauses after "thought" and idioms ("made sense",
"looked after") are never raised, and shown instances are spread across the
book rather than taken from its opening. prose-norms sets the finding aside
in explanation, speech and reflection, with a reason, and keeps it in action
and description.
Repetition: a capitalised word whose lowercase form never occurs in the text
is a name, and names are not repetitions.
The editor now honours what the engine decided: findings set aside for their
passage render muted with the reason in the tooltip and no fix button;
per-type counts and category lists are of the findings that actually scored.
A synonym from a fixed list is never auto-applied for the author (the button
becomes Edit Here), and the fixer no longer picks a replacement at random.
Also fixes the double-escaped "Replace & Fix" label.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VWzYm8L13Mra3ihjhuCM2U
…emoved
Characters: analyzeCharacters built a sentence-start set and never used it,
so every capitalised word with three mentions was a character ("Your",
"Every", "Consider"). A word is now a name only where grammar did not force
the capital: it must not occur in lowercase anywhere in the text, and must
be capitalised mid-sentence at least twice (or carry a speech attribution).
No list of names is involved. The card is titled "Names in the text" for
nonfiction.
Lexical diversity: a raw unique/total ratio falls with length by arithmetic
(a 70,000-word book scores ~9% however varied), so it is now a moving-average
type/token ratio over 500-word windows, stable across length, and labelled as
such beside the raw unique-word count. Score bands re-based for that scale.
Pacing heatmap: the 200-word segments are now coloured by the five-mode
passage classifier already used for context, taking the mode that covers most
of each segment and calling ambiguous stretches mixed. The 27-verb keyword
estimate remains only as a fallback and is labelled when used.
Line editing: every sub-score was base + bonuses that topped out between 45
and 70, so a flawless manuscript could not score above ~61 on Line Editing.
Sub-scores now sum to 100. Uncapped raw-count penalties (tone clash, long
dialogue paragraphs, filter phrases) are now rates per 1,000 words or shares,
capped. A quoted term no longer counts as a dialogue paragraph. Prose with no
dialogue no longer forfeits the dialogue components. POV is N/A for
nonfiction and left out of the composite rather than scored a fake 45; the
nonfiction flow handling lives inside the function so sub-scores and
composite always agree.
Also brings analyzer-worker.js's import versions in step with app.html; it
was still loading analyzer.js?v=25.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VWzYm8L13Mra3ihjhuCM2U
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a6fece545c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (verb === 'got' && /\b(have|has|had|'ve|'s|'d|’ve|’s|’d)\s+$/i.test(before)) return true; | ||
| const afterTrim = text.slice(index + verb.length, index + verb.length + 40).replace(/^\s+/, ''); | ||
| if (!afterTrim) return false; | ||
| if (this.WEAK_VERB_PARTICLES.test(afterTrim)) return true; |
There was a problem hiding this comment.
Keep directional verb uses in the reliance count
When a movement verb is followed by an ordinary direction or preposition, such as walked up the hill, walked down the lane, or walked through the gate, this blanket particle check classifies the use as idiomatic and removes it before the minimum-use and reliance calculations. A manuscript relying entirely on these common constructions can therefore receive no weak-verb finding; restrict exclusions to actual verb-specific phrasal idioms.
Useful? React with 👍 / 👎.
| (raw.match(/\b[A-Z][a-z]{3,}\b/g) || []).forEach(tok => { | ||
| const w = tok.toLowerCase(); | ||
| if (!lowercaseForms.has(w)) names.add(w); |
There was a problem hiding this comment.
Distinguish sentence openers from character names
When a repeated common word appears only at sentence starts, such as Darkness covered the road. Darkness swallowed the trees., both capitalized occurrences are added to names because no lowercase form exists, so the repetition is silently suppressed. Capitalization at a sentence boundary is not evidence of a name; exclude sentence-initial tokens from this name heuristic or reuse the positional logic from analyzeCharacters.
Useful? React with 👍 / 👎.
| const issuesPerK = Math.round(allIssues.length / Math.max(totalWords, 1) * 1000 * 10) / 10; | ||
| // Issue density normalized per 1000 words, over the findings that actually scored. | ||
| const issuesPerK = Math.round(scoredIssues.length / Math.max(totalWords, 1) * 1000 * 10) / 10; | ||
| const countScored = t => scoredIssues.reduce((n, i) => n + (i.type === t ? 1 : 0), 0); |
There was a problem hiding this comment.
Count only narrative findings beside narrative scores
For documents with segmented front or back matter, scoreCopyEditing is computed from narrativeIssues, but countScored uses every non-context-suppressed issue in the full document. Consequently, issues in copyright pages, dedications, or appendices appear beside a score they did not affect, contradicting the new count/score consistency guarantee; derive these counts from the same narrative-filtered collection used for scoring.
Useful? React with 👍 / 👎.
…e manuscript-derived metrics
Both branches repaired the same Detailed-tab numbers from the same screenshots.
Resolution, dimension by dimension:
- Line editing: upstream's subtract-from-100 model with capped, length-normalised
penalties is taken whole. Two fixes are ported into it: a dialogue paragraph must
carry a spoken line (capital after the quote, punctuation before the close), so a
quoted term no longer counts; and sentence continuity is N/A for nonfiction rather
than judged on word overlap, with the composite taken over applicable dimensions.
- Style: upstream's MSTTR-100 lexical diversity and its removal of POV/word-length
as quality points are taken; the MATTR helper is dropped.
- Characters: the manuscript-derived rules are kept (a word that occurs in lowercase
anywhere is not a name; a name must be capitalised mid-sentence or carry person
evidence). Upstream's skip list, which named the words from the screenshot
("Patience Poverty Take Start … Mark Gates Nigeria"), is not taken: it would hide
a character called Mark in the next book. Upstream's genre gate (expository
nonfiction: not applicable), two-word names and possessive evidence are kept.
- Heatmap: the passage-classifier segments are kept and now carry upstream's
provenance (word range, excerpt, counts); upstream's improved keyword estimate is
the fallback and is labelled when used.
- Issue counts: counts of findings that scored, excluding those set aside for their
passage, so the number beside a score is the number that produced it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VWzYm8L13Mra3ihjhuCM2U
No description provided.