Found while working on #116.
What happens. The two table parsers disagree about what counts as a table row. Excel normalises a row that is missing its trailing |; Word requires both pipes and stops the table at that line. The same markdown produces a different table in each tool.
Reproduced against b36ebfb:
markdown:
'| A | B |'
'|---|---|'
'| 1 | 2 |'
'| 3 | 4' <- trailing pipe missing
Excel rows : [['A', 'B'], ['1', '2'], ['3', '4']]
Word rows : [['A', 'B'], ['1', '2']]
Word does not merely drop the row. parse_table() returns next_idx=3, pointing at the offending line, so the block dispatcher picks it up as ordinary text and the document renders a literal paragraph reading | 3 | 4 underneath the table. The user sees pipe characters in their document, which reads as a rendering bug rather than as malformed input.
Why. xlsx_tools/helpers.py:273-275:
if line.startswith('|'):
# Normalize: ensure trailing pipe for consistent splitting
if not line.endswith('|'):
docx_tools/block_elements.py:91:
if line.startswith('|') and line.endswith('|'):
Which one is right. Excel's, in my view. A trailing pipe is optional in GitHub-flavoured Markdown — | 3 | 4 is a valid row and most editors render it as one — and the tool's whole premise is that a model writes the markdown, where a dropped trailing pipe is a plausible slip. Being lenient and consistent beats being strict in one tool only. Worth an explicit decision either way, since the alternative (make Excel strict too) is defensible and at least ends the disagreement.
Fix. Give Word the same normalisation Excel has, at block_elements.py:91. Note that _is_separator_line() was extracted in #116 precisely so the guard and the loop could not drift apart on separator rows; this is the same class of drift across two packages rather than within one, so consider whether the row-detection predicate belongs somewhere both parsers can share rather than being fixed twice.
Warning channel. Whichever way it is decided, the rejecting parser should report it: a row the caller wrote and the document does not contain is exactly what docx_tools/warnings.py is for (#114), and by the AGENTS.md rubric it is warning severity — the content is there but not as asked. Today it is silent apart from the stray paragraph.
Tests. One case per parser asserting they agree on the same input, confirmed to fail on the commit before the fix. tests/test_docx_base.py and the xlsx table tests are the natural homes; a shared case would be better if the predicate ends up shared.
Found while working on #116.
What happens. The two table parsers disagree about what counts as a table row. Excel normalises a row that is missing its trailing
|; Word requires both pipes and stops the table at that line. The same markdown produces a different table in each tool.Reproduced against
b36ebfb:Word does not merely drop the row.
parse_table()returnsnext_idx=3, pointing at the offending line, so the block dispatcher picks it up as ordinary text and the document renders a literal paragraph reading| 3 | 4underneath the table. The user sees pipe characters in their document, which reads as a rendering bug rather than as malformed input.Why.
xlsx_tools/helpers.py:273-275:docx_tools/block_elements.py:91:Which one is right. Excel's, in my view. A trailing pipe is optional in GitHub-flavoured Markdown —
| 3 | 4is a valid row and most editors render it as one — and the tool's whole premise is that a model writes the markdown, where a dropped trailing pipe is a plausible slip. Being lenient and consistent beats being strict in one tool only. Worth an explicit decision either way, since the alternative (make Excel strict too) is defensible and at least ends the disagreement.Fix. Give Word the same normalisation Excel has, at
block_elements.py:91. Note that_is_separator_line()was extracted in #116 precisely so the guard and the loop could not drift apart on separator rows; this is the same class of drift across two packages rather than within one, so consider whether the row-detection predicate belongs somewhere both parsers can share rather than being fixed twice.Warning channel. Whichever way it is decided, the rejecting parser should report it: a row the caller wrote and the document does not contain is exactly what
docx_tools/warnings.pyis for (#114), and by the AGENTS.md rubric it iswarningseverity — the content is there but not as asked. Today it is silent apart from the stray paragraph.Tests. One case per parser asserting they agree on the same input, confirmed to fail on the commit before the fix.
tests/test_docx_base.pyand the xlsx table tests are the natural homes; a shared case would be better if the predicate ends up shared.