A small, self-contained Python project that converts a Markdown (.md) file
into a faithful .docx (Word) document. Its focus is conversion fidelity.
The output is a real Word document that visually and structurally matches the
Markdown source as rendered by GitHub / CommonMark.
The implementation is a token-tree walker (via mistune v3), not string
replacement. Formatting nests correctly, whitespace in code is preserved, and
edge cases (links inside bold, ordered lists with custom start numbers, **
inside fenced code) round-trip faithfully.
- Python 3.10+
mistune(v3): Markdown parser (AST)python-docx: DOCX writer
Install with:
pip install -r requirements.txtRun the tests:
pip install pytest # dev-only
python -m pytestpython md2docx.py input.md -o output.docx
python md2docx.py input.md # writes input.docx
python md2docx.py --help
| Flag | Description |
|---|---|
-o, --output PATH |
Output .docx path (default: <input>.docx) |
--font NAME |
Base font family (default: Calibri) |
--font-size PT |
Base font size in points (default: 11) |
--template PATH.docx |
Optional Word template/theme to base styles on |
--no-remote-images |
Skip downloading remote images; use alt text instead |
from md2docx import convert
from md2docx.converter import ConversionOptions
text = open("README.md", encoding="utf-8").read()
doc = convert(text, out_path="README.docx",
options=ConversionOptions(font="Aptos", font_size=12))| Path | Purpose |
|---|---|
md2docx/model.py |
Parser-independent intermediate document model |
md2docx/parser.py |
Mistune AST → model (swappable parser) |
md2docx/writer.py |
Model → .docx via python-docx (Word-generation) |
md2docx/converter.py |
Orchestration: parse → resolve images → render |
md2docx/cli.py |
CLI entry point |
md2docx.py |
Thin script wrapper for python md2docx.py ... |
tests/ |
pytest suite + kitchen-sink fixture |
examples/ |
Checked-in .md → .docx sample pair |
The parser and the writer are decoupled by the model. Swapping the Markdown parser for another that emits the same model never touches the Word-generation code.
| Markdown | DOCX output |
|---|---|
**bold**, __bold__ |
Bold run |
*italic*, _italic_ |
Italic run |
***bold italic*** |
Bold + italic run |
~~strikethrough~~ |
Strikethrough run |
`inline code` |
Monospace (Consolas) run with light gray shading |
[text](url) |
Real clickable hyperlink (<w:hyperlink>) |
# … ###### |
Built-in Heading 1-6 styles (outline/TOC friendly) |
| paragraphs | Normal paragraphs; soft/hard breaks → <w:br/> |
> quote / >> nested |
Word "Quote" style with increasing indentation |
```lang |
Shaded, bordered monospace paragraphs; preserve lines |
---, ***, ___ |
Horizontal rule (bottom-bordered paragraph) |
-/*/+ lists |
Real Word bulleted numbering |
1.… lists |
Real Word decimal numbering (custom start respected) |
| nested/mixed lists | Correct Word indentation + per-list numbering |
- [ ] / - [x] |
Task list with ☐ / ☑ prefixes |
| pipe tables | Native Word tables, header bold + shading, alignment |
 local |
Embedded image, scaled down to page width if oversized |
 |
Downloaded and embedded (disable with --no-remote-images) |
[^n] footnotes |
End-note style numbered "Notes" section at document end |
| YAML front matter | title/author/date → Word document properties |
- Word footnotes are not produced. python-docx does not expose Word's
native footnote storage, so
[^n]references are rendered as an end-note style numbered "Notes" section at the end of the document. Inline footnote markers appear as[1],[2], … superscripts. - Footnote body formatting is flattened to plain text. Rich inline formatting inside a footnote definition is converted to plain text (the markdown markers are stripped, but bold/italic is not carried over).
- Task lists use Unicode glyphs ☐ / ☑. If the active font has no glyph for these, Word falls back to another font; on very limited fonts they may render as boxes instead.
- Light-weight YAML front matter. Only scalar
key: valuelines are read (specificallytitle,author,date). Lists/maps/nested YAML in front matter are skipped, and front matter that spans an unclosed---is treated as body text. - Remote images require network access at conversion time. If a remote
image fails to download (offline, timeout, error page), the alt text is
shown in brackets instead. Use
--no-remote-imagesto never attempt downloads. - Tables use an even column layout. python-docx tables default to auto-fit; column widths are not read from the Markdown source (Markdown has no width information anyway).
- No syntax highlighting. Fenced code block language tags are preserved in the intermediate model but not rendered as colored syntax.
<kbd>/HTML passthrough is not converted. Any inline HTML in the Markdown flows through as raw text inmistuneand is not rendered as HTML elements in Word.- Numbering visuals depend on Word's list rendering. Bullet glyph indentation is set per level; deeply nested lists beyond the built-in indentation may need manual tweaks in Word.
- Soft breaks render as line breaks. Plain newlines inside a paragraph
become hard line breaks (
<w:br/>) rather than being collapsed to spaces. This keeps the source line structure intact, which most people expect, but differs from pure CommonMark HTML rendering, where a bare newline collapses to a space.
The pytest suite (in tests/) includes:
- a kitchen-sink fixture exercising every feature above,
- assertions on the resulting
.docxXML structure via python-docx's object model (run.bold/.italic, paragraph.style.name, table cell contents, hyperlink targets, numbering definitions), - specific regression cases covering the trickier edge cases: nested bold/italic, links inside bold, formatting inside table cells, ordered lists with a custom start number, and code blocks whose content looks like Markdown syntax.