Skip to content

About

A deterministic prose linter for formulaic constructions, and a review chain where two model families propose rewrites that only survive if the linter agrees. Invented text, 50-sentence labelled corpus, 29 tests, standard library only.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Prose check chain

A cover letter that reads as formulaic gets skimmed. The constructions that make it read that way are surprisingly few and surprisingly regular: the em dash mid-sentence, "not X, but Y", "what carries over is", the stock phrases, three short nouns in a row, five sentences of the same length. This project does two things with that observation.

  1. A deterministic linter finds those constructions with regular expressions and two sentence statistics. Same text, same report, no model involved.
  2. A review chain hands the text and the linter's findings to two reviewers in turn, a "cold reader" and a model from a different vendor, and applies their rewrites one at a time. A rewrite survives only if the linter, run again, agrees that it did not make things worse.

The point of the second part is not the models. It is that the models never get the last word: the deterministic pass judges every suggestion, and the result is a number (findings before, findings after, rounds, rejected suggestions) rather than an impression.

Everything in this repository is invented. The example letters, the corpus sentences, the company and the applicant do not exist.

Run it

python -m venv .venv
.venv/Scripts/activate        # Windows
# source .venv/bin/activate   # Linux, macOS
pip install -e ".[dev]"

python -m prosecheck lint examples/letter_en.md --lang en   # the deterministic pass
python -m prosecheck score                                  # the linter against its corpus
python -m prosecheck chain examples/letter_en.md --lang en  # the chain, replaying stored answers
pytest                                                      # 29 tests

lint exits 1 when a FAIL is left, chain exits 1 when the final text still has one. Both print what they found and why.

What the linter finds

rule level what it is
em-dash, spaced-en-dash FAIL a dash used as punctuation; ranges like 30–40 % are allowed
placeholder FAIL [TODO], [...] and friends left in the text
de-antithesis, en-antithesis FAIL "nicht X, sondern Y", "not only X but Y"; a concessive "not yet, but I" is allowed
de-transition, en-trigger FAIL "darüber hinaus", "furthermore", "leverage", "cutting-edge" and the rest of the list
de-pointer, en-pointer WARN "genau das", "exactly the kind": points back instead of saying something
de-dramatic-colon, en-dramatic-colon WARN announcement, colon, lower-case reveal
de-triplet, en-triplet WARN "Python, SQL und Git": three short words in a row
en-contrast, en-pseudo-cleft WARN "rather than", "what carries over is"
fragment-chain WARN three sentences of seven words or fewer in a row
uniform-length WARN five or more sentences with a standard deviation under four words

FAIL means the chain will not pass until it is gone. WARN is left to a reader or a reviewer. Label lines ("Skills: Python, SQL and Git") are recognised and exempt from the colon and triplet rules, because a colon in a header is a separator, not a pause.

The list is one person's list, kept from the letters that person writes and reads. It has no basis in the literature and does not claim one. What it has is a corpus.

The corpus and the score

src/prosecheck/data/corpus.jsonl holds 50 invented sentences, each labelled with the rule ids a reader expects, or with none. python -m prosecheck score runs the linter over them:

corpus entries 50, clean 18, clean entries with a finding 0
expected findings 36, raised 36 of them, unexpected findings 0

Every rule has at least one positive example and the clean entries include the traps: number ranges with an en dash, "M.Sc." not splitting a sentence, a header line with a colon, a concessive "but I", a single short sentence that is fine on its own. The test suite fails if the linter and the corpus ever disagree.

A perfect score on a corpus written by the same hand that wrote the rules is not evidence of much. It is evidence that the rules do what they say, which is the property the chain depends on.

The chain

lint ─► FAIL/WARN? ─► cold reader ─► apply ─► lint
                  └─► cross family ─► apply ─► lint

Both reviewers get the same brief (review.py): the text, the house rules, the linter's findings, and an instruction to propose the smallest rewrite that removes each construction without adding claims. They answer in one JSON shape, {"quote", "issue", "rewrite", "rule"}. Then three things happen to every suggestion, in chain.py:

  • If quote does not occur exactly once in the text, it is dropped and counted. Nothing is guessed at.
  • The rewrite is applied on its own and the text is linted again. If the FAIL count went up, the rewrite is rejected and counted. Trading an antithesis for an em dash does not pass.
  • Otherwise it stays.

The loop stops when the lint is clean, when a round applied nothing, or at the round limit (default 3). The chain never edits text itself.

Two reviewers from two vendors because one model family shares one set of habits. A model that writes "furthermore" will not reliably see "furthermore".

The example, measured

examples/letter_en.md opens with six FAIL and four WARN findings. Through the chain, with the stored answers:

FAIL 6 -> 0, WARN 4 -> 0, rounds 2, applied 5, rejected 0, dropped 0

The German example ends with FAIL 3 -> 0 and one warning left: uniform-length. The reviewers shortened every sentence into the same rhythm, and the linter says so. The chain reports it rather than hiding it. That warning is the most useful line in the output: a text that passes every rule can still read as built.

About the stored answers

The files under fixtures/ were authored for this repository, not recorded from a live model call. They went through the same recording path a live provider uses, so replay and record are exercised, but they were written during development (with the same AI assistance as the rest of the code) to show the chain working. They say nothing about how a given model would actually answer. Their provider field says so.

To record real answers, set the keys and run once with --provider live --record:

export ANTHROPIC_API_KEY=...  ANTHROPIC_MODEL=...
export OPENAI_API_KEY=...     OPENAI_MODEL=...
python -m prosecheck chain examples/letter_en.md --lang en --provider live --record

No default model is assumed; both have to be named. The HTTP client is urllib from the standard library, one request per review, no retries.

What this does not show

  • Detection. This is not a detector for machine-written text and does not become one by inversion. It removes a list of constructions; a text can be free of all of them and still be empty.
  • Judgement about content. Whether a claim in the letter is true, or whether the letter fits the job, is outside the rules. A reviewer with the job advertisement in front of it is a different tool.
  • Real reviewer behaviour. The shipped fixtures are stand-ins; the numbers above describe the mechanism, not a model.
  • German capitalisation. The dramatic colon rule only sees a lower-case word after the colon. "Das ist der Punkt: Prüfen kostet Zeit" slips through. The corpus records this as a known gap.

Layout

src/prosecheck/
  rules.py       the catalogue: one regex, level, language and reason per rule
  lint.py        the deterministic pass and the two sentence statistics
  corpus.py      the labelled corpus and the score against it
  review.py      reviewer brief, JSON shape, parsing and validation
  providers.py   replay, record, one plain HTTPS call
  chain.py       lint, review, apply, lint again
  cli.py         lint / score / chain
  data/corpus.jsonl
examples/        two invented letters, DE and EN
fixtures/        stand-in reviewer answers for replay
tests/
  test_lint.py      one test per rule family and per trap
  test_corpus.py    linter and corpus must agree
  test_chain.py     parsing, rejection, stop conditions, record and replay
  test_examples.py  the two letters end to end

On AI assistance

This repository was written with AI assistance. The rule list, the decision to let the linter judge the reviewers, and the review of the result are mine. The stand-in reviewer answers in fixtures/ came out of the same assisted work; they are labelled as stand-ins wherever they are used.

Licence

MIT. See LICENSE.

About

A deterministic prose linter for formulaic constructions, and a review chain where two model families propose rewrites that only survive if the linter agrees. Invented text, 50-sentence labelled corpus, 29 tests, standard library only.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages