feat(sitegen): group thousands with U+202F, and route the last renderer through numbers - #18
Merged
Merged
Conversation
…er through numbers
`0007` §5 clause 8 of the portfolio page specification asks every grouped figure
on a published page for a narrow no-break space. This repository wrote three, in
three spellings from three places, which is the exact drift `sitegen/numbers.py`
was written to end.
`numbers.integer()` now groups with U+202F, written `"\u202f"` rather than as the
character: U+0020 and U+202F are one string in a diff, a terminal and a `grep`,
and a module whose whole purpose is that surfaces cannot disagree about a number
should not depend on a reader spotting an invisible one.
The other two were not reached by changing that function, and the interesting one
is the second:
- `sitegen/page.py:219` and its hand-typed twin in `README.md` carry `1 000-resample`
in prose. Both move, so the two surfaces keep saying the same sentence.
- `examples/validation_table.py:130` had a format string of its own — `{:,}` — in
direct contradiction of `numbers.py`'s opening line, "no renderer is allowed a
format string of its own", and of `docs/decisions/0007-the-page-is-generated.md`,
which tabulates this exact script as the divergence the package argues against.
It now calls `numbers.integer()`. The number it writes, `n = 14 745/arm`, reaches
`docs/data/findings.json`, the page and the README verbatim.
`tests/test_record.py:86` re-derives that figure closed-form and asserted it with
`{:,}`. It now applies the substitution itself rather than calling `numbers.integer`
— a guard that formats through the function it guards asserts nothing about the
formatting.
Re-recording `docs/data/findings.json` moved two fields beyond the scenario string,
and both are worth stating. `recorded_on` is today. `ab_lab_version` was `0.3.0.dev0`
and is now `0.4.1`: the published provenance named a version two minor releases
behind the package. **Every one of the five measured rates reproduced to the digit**
— 0.0538, 0.0497, 0.8077, 0.7929, 0.0110 — so the seed is honest and the version
gap moved no result.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
0007§5 clause 8 of the portfolio page specification asks every grouped figure on apublished page for a narrow no-break space, U+202F. This repository writes three figures, in
three spellings from three places — which is the drift
sitegen/numbers.pyexists to end,reproduced inside the repository that argues against it.
10,000numbers.integer()1 000-resamplesitegen/page.py:219, and a hand-typed twin inREADME.md14,745examples/validation_table.py:130, a format string of its ownThe third one is the point
sitegen/numbers.pyopens with "no renderer is allowed a format string of its own", anddocs/decisions/0007-the-page-is-generated.md:19-22tabulates this exact script as thedivergence the package argues against. It still had
{:,}. It now callsnumbers.integer(),and the figure it writes —
n = 14 745/arm— reachesdocs/data/findings.json,docs/index.htmlandREADME.mdverbatim.The separator is written
"\u202f"in Python sources rather than as the character. U+0020 andU+202F are one string in a diff, a terminal and a
grep, and a module whose whole purpose isthat surfaces cannot disagree about a number should not rely on a reader spotting an invisible
one.
README.mdgets the character, being Markdown.What the re-record turned up
--recordmoved two fields beyond the scenario string:recorded_on:2026-08-21→2026-09-08.ab_lab_version:0.3.0.dev0→0.4.1. The published provenance named a version twominor releases behind the package.
Every one of the five measured rates reproduced to the digit — 0.0538, 0.0497, 0.8077,
0.7929, 0.0110 — so the seed is honest and the version gap moved no result. That is the check
worth having, and it is why the re-record is reported here rather than assumed.
What
tests/test_record.py:86actually guardsIt re-derives the sample size closed-form and asserts the recorded scenario names it. It is
re-pinned to the new spelling, and it applies the substitution itself rather than calling
numbers.integer— a guard that formats through the function it guards asserts nothing aboutthe formatting.
Mutation-tested, and the result corrects a natural assumption about it:
validation_table.py:130back to{:,}, without re-recordingdocs/data/findings.jsonback to a commavalidation_table.py:130back to{:,}and re-recordednumbers.integerback to a commatest_site_committed.py's byte guardsSo the guard is live over the recorded evidence and over the script through a re-record, and
is deliberately blind to an un-recorded script edit. That is correct —
findings.jsonis acommitted artifact — but it is not what the test's name suggests, so it is written down.
Verification
pytest— full suite green.python examples/validation_table.py --record, thenpython -m sitegen.build, whichrewrote
docs/index.htmland thegenerated:regions ofREADME.md.clear:ok 8 separator U+202F 3.