A cover letter that reads as formulaic gets skimmed. The constructions that make it read that way are surprisingly few and surprisingly regular: the em dash mid-sentence, "not X, but Y", "what carries over is", the stock phrases, three short nouns in a row, five sentences of the same length. This project does two things with that observation.
- A deterministic linter finds those constructions with regular expressions and two sentence statistics. Same text, same report, no model involved.
- A review chain hands the text and the linter's findings to two reviewers in turn, a "cold reader" and a model from a different vendor, and applies their rewrites one at a time. A rewrite survives only if the linter, run again, agrees that it did not make things worse.
The point of the second part is not the models. It is that the models never get the last word: the deterministic pass judges every suggestion, and the result is a number (findings before, findings after, rounds, rejected suggestions) rather than an impression.
Everything in this repository is invented. The example letters, the corpus sentences, the company and the applicant do not exist.
python -m venv .venv
.venv/Scripts/activate # Windows
# source .venv/bin/activate # Linux, macOS
pip install -e ".[dev]"
python -m prosecheck lint examples/letter_en.md --lang en # the deterministic pass
python -m prosecheck score # the linter against its corpus
python -m prosecheck chain examples/letter_en.md --lang en # the chain, replaying stored answers
pytest # 29 testslint exits 1 when a FAIL is left, chain exits 1 when the final text still
has one. Both print what they found and why.
| rule | level | what it is |
|---|---|---|
em-dash, spaced-en-dash |
FAIL | a dash used as punctuation; ranges like 30–40 % are allowed |
placeholder |
FAIL | [TODO], [...] and friends left in the text |
de-antithesis, en-antithesis |
FAIL | "nicht X, sondern Y", "not only X but Y"; a concessive "not yet, but I" is allowed |
de-transition, en-trigger |
FAIL | "darüber hinaus", "furthermore", "leverage", "cutting-edge" and the rest of the list |
de-pointer, en-pointer |
WARN | "genau das", "exactly the kind": points back instead of saying something |
de-dramatic-colon, en-dramatic-colon |
WARN | announcement, colon, lower-case reveal |
de-triplet, en-triplet |
WARN | "Python, SQL und Git": three short words in a row |
en-contrast, en-pseudo-cleft |
WARN | "rather than", "what carries over is" |
fragment-chain |
WARN | three sentences of seven words or fewer in a row |
uniform-length |
WARN | five or more sentences with a standard deviation under four words |
FAIL means the chain will not pass until it is gone. WARN is left to a reader or a reviewer. Label lines ("Skills: Python, SQL and Git") are recognised and exempt from the colon and triplet rules, because a colon in a header is a separator, not a pause.
The list is one person's list, kept from the letters that person writes and reads. It has no basis in the literature and does not claim one. What it has is a corpus.
src/prosecheck/data/corpus.jsonl holds 50 invented sentences, each labelled
with the rule ids a reader expects, or with none. python -m prosecheck score
runs the linter over them:
corpus entries 50, clean 18, clean entries with a finding 0
expected findings 36, raised 36 of them, unexpected findings 0
Every rule has at least one positive example and the clean entries include the traps: number ranges with an en dash, "M.Sc." not splitting a sentence, a header line with a colon, a concessive "but I", a single short sentence that is fine on its own. The test suite fails if the linter and the corpus ever disagree.
A perfect score on a corpus written by the same hand that wrote the rules is not evidence of much. It is evidence that the rules do what they say, which is the property the chain depends on.
lint ─► FAIL/WARN? ─► cold reader ─► apply ─► lint
└─► cross family ─► apply ─► lint
Both reviewers get the same brief (review.py): the text, the house rules,
the linter's findings, and an instruction to propose the smallest rewrite
that removes each construction without adding claims. They answer in one
JSON shape, {"quote", "issue", "rewrite", "rule"}. Then three things
happen to every suggestion, in chain.py:
- If
quotedoes not occur exactly once in the text, it is dropped and counted. Nothing is guessed at. - The rewrite is applied on its own and the text is linted again. If the FAIL count went up, the rewrite is rejected and counted. Trading an antithesis for an em dash does not pass.
- Otherwise it stays.
The loop stops when the lint is clean, when a round applied nothing, or at the round limit (default 3). The chain never edits text itself.
Two reviewers from two vendors because one model family shares one set of habits. A model that writes "furthermore" will not reliably see "furthermore".
examples/letter_en.md opens with six FAIL and four WARN findings. Through
the chain, with the stored answers:
FAIL 6 -> 0, WARN 4 -> 0, rounds 2, applied 5, rejected 0, dropped 0
The German example ends with FAIL 3 -> 0 and one warning left:
uniform-length. The reviewers shortened every sentence into the same
rhythm, and the linter says so. The chain reports it rather than hiding it.
That warning is the most useful line in the output: a text that passes every
rule can still read as built.
The files under fixtures/ were authored for this repository, not
recorded from a live model call. They went through the same recording path
a live provider uses, so replay and record are exercised, but they were
written during development (with the same AI assistance as the rest of the
code) to show the chain working. They say nothing about how a given model
would actually answer. Their provider field says so.
To record real answers, set the keys and run once with --provider live --record:
export ANTHROPIC_API_KEY=... ANTHROPIC_MODEL=...
export OPENAI_API_KEY=... OPENAI_MODEL=...
python -m prosecheck chain examples/letter_en.md --lang en --provider live --recordNo default model is assumed; both have to be named. The HTTP client is
urllib from the standard library, one request per review, no retries.
- Detection. This is not a detector for machine-written text and does not become one by inversion. It removes a list of constructions; a text can be free of all of them and still be empty.
- Judgement about content. Whether a claim in the letter is true, or whether the letter fits the job, is outside the rules. A reviewer with the job advertisement in front of it is a different tool.
- Real reviewer behaviour. The shipped fixtures are stand-ins; the numbers above describe the mechanism, not a model.
- German capitalisation. The dramatic colon rule only sees a lower-case word after the colon. "Das ist der Punkt: Prüfen kostet Zeit" slips through. The corpus records this as a known gap.
src/prosecheck/
rules.py the catalogue: one regex, level, language and reason per rule
lint.py the deterministic pass and the two sentence statistics
corpus.py the labelled corpus and the score against it
review.py reviewer brief, JSON shape, parsing and validation
providers.py replay, record, one plain HTTPS call
chain.py lint, review, apply, lint again
cli.py lint / score / chain
data/corpus.jsonl
examples/ two invented letters, DE and EN
fixtures/ stand-in reviewer answers for replay
tests/
test_lint.py one test per rule family and per trap
test_corpus.py linter and corpus must agree
test_chain.py parsing, rejection, stop conditions, record and replay
test_examples.py the two letters end to end
This repository was written with AI assistance. The rule list, the decision
to let the linter judge the reviewers, and the review of the result are
mine. The stand-in reviewer answers in fixtures/ came out of the same
assisted work; they are labelled as stand-ins wherever they are used.
MIT. See LICENSE.