A framework for extracting, encoding, and reproducing any author's writing voice using AI.
Voice DNA is a structured methodology for analyzing a body of writing and producing a detailed, machine-readable voice profile. The profile can then be used as a system prompt to generate new content that sounds like the original author -- not a surface-level imitation, but a deep reproduction of their rhythm, word choice, punctuation habits, rhetorical patterns, and thinking style.
app-script-gmail-extraction.md -- A step-by-step guide to exporting your sent emails from Gmail into a Google Doc using Google Apps Script. This is a corpus-building tool -- if you don't already have a body of writing to analyze, your sent emails are one of the richest sources of your natural voice. The guide walks through setup, configuration, running the script, and cleaning the output.
voice-dna-extraction-rulebook.md -- The step-by-step process for building a Voice DNA Profile from scratch. Covers five phases:
- Framework -- Scope the analysis model to the corpus
- Batch Analysis -- Read the corpus in batches, extract observations across 7 linguistic layers (orthography, punctuation, lexicon, syntax, cadence, discourse strategy, epistemic posture)
- Synthesis -- Merge observations into the 8 profile deliverables
- Validation -- Generate test content, score it, diagnose failures
- Raw Material Supplement -- Expand the profile with unedited/spoken material to capture the full register range
voice-dna-profile-template.md -- A blank profile with every section, field, and table ready to fill in. Includes guidance prompts at each field explaining what to look for. Sections:
- Voice DNA Card -- One-page cheat sheet (10 fields: tone, formality, persona, devices, phrases, lexicon, grammar, sentence style, perspective, rhetorical moves)
- Voice Constitution -- 50-80 concrete do/don't rules organized by linguistic layer, with signal rankings and evidence annotations
- Reference Tables -- Synonym preferences, transition preferences, per-1000-word frequency targets
- Sentence Skeleton Library -- 30-50 abstract sentence patterns extracted from the corpus, organized by function
- Mode Router -- Primary writing mode + variants, with a feature deltas table and intra-piece routing rules
- Production Rewrite Prompt -- A self-contained system prompt with Anti-Caricature Governor, Anti-Polish Directive, Content Lock, Belief Leakage Prevention, and blacklists
- Evaluation Rubric -- 22-item yes/no checklist for scoring voice match
- Iteration Log -- Tracking adjustments over time
- Collect 20,000+ words of writing from your target author
- Open the rulebook and follow it phase by phase
- Fill in the profile template as you go (Phase 3)
- Validate the profile (Phase 4)
- Paste the Production Rewrite Prompt into any AI conversation to write in the extracted voice
The process works with any AI assistant that can read and analyze text. It was built with Claude but is model-agnostic.
The framework analyzes writing across seven layers, from surface mechanics to deep cognitive posture:
| Layer | Name | What it captures |
|---|---|---|
| L1 | Orthography & Visual Style | Spelling, capitalization, emphasis formatting |
| L2 | Punctuation & Micro-Prosody | Punctuation as rhythm and tone |
| L3 | Lexicon & Collocations | Word choice, vocabulary range, recurring phrases |
| L4 | Syntax & Sentence Geometry | Sentence structure, length, complexity |
| L5 | Cadence & Paragraph Choreography | Paragraph patterns, rhythm, pacing |
| L6 | Discourse Strategy & Rhetorical Moves | Argument structure, openings, closings, devices |
| L7 | Epistemic Posture & Thinking Style | Confidence, authority, reader relationship |
Negative Stylometry -- What the author never does is as diagnostic as what they always do. If someone never uses semicolons, a generated text with semicolons fails regardless of what else it gets right.
Anti-Caricature Governor -- Real voice is unevenly distributed. Some paragraphs are plain, some are loaded with signature features. Concentrating every voice marker into every paragraph produces a caricature, not a reproduction.
Anti-Polish Directive -- Sentence fragments, comma splices, unconventional grammar, and casual spelling are often voice features, not errors. The directive prevents AI from "correcting" them during generation.
Mode Router -- Most authors don't have one voice. They have a primary mode and several variants (teaching vs. storytelling vs. ranting vs. reflective). The Mode Router maps when each activates and how the features shift.
Content Lock -- Voice is applied to content, not instead of content. All source facts must survive voice transformation. No dropping claims for stylistic reasons.
| Corpus Size | Profile Quality |
|---|---|
| 5,000-10,000 words | Thin -- enough for a Voice DNA Card and basic Constitution, but patterns may not be stable |
| 10,000-20,000 words | Workable -- enough for most deliverables, but some rules will be under-evidenced |
| 20,000-40,000 words | Strong -- enough to fill the full profile with well-evidenced rules |
| 40,000+ words | Robust -- pattern stability is high, mode variants become visible, frequency targets are reliable |
| + Raw material | Expanded -- captures the full register range from performing voice to thinking voice |
MIT