Summary
The multi_tolerance scorer in src/evaluation/reward.py assumes English/US number formatting. On non-English answers this does not fail loudly — it silently scores equivalent values as wrong, and in one case silently rescales a number by 10x. Since _score_multi_tolerance is the default scorer and its output drives frontier selection, an evolution run on a non-English task set optimises against a largely meaningless signal.
Filing this as a report rather than a PR because the fix involves a design choice (how to decide a locale) that is yours to make — see "Open question" at the end.
Environment
main @ 36f6f04
- Python 3.12, Linux
- No configuration; default scorer path
Reproduction
Save as repro_locale.py in the repo root and run python repro_locale.py:
from src.evaluation.reward import (
score_answer,
extract_numbers_with_context,
detect_unit_in_context,
)
print("1) decimal comma is silently rescaled")
print(" extract_numbers_with_context('4,2') ->", extract_numbers_with_context('4,2'))
print("2) space thousands separator")
print(" score_answer('2602', '2 602') ->", score_answer('2602', '2 602'))
print(" score_answer('2602', '2,602') ->", score_answer('2602', '2,602'))
print("3) non-English scale words")
for c in ('4.2 billion', '4.2 миллиарда', '543 миллиона'):
print(f" detect_unit_in_context({c!r}) -> {detect_unit_in_context(c)}")
print("4) end-to-end")
for gt, pred in (('4.2', '4,2'), ('1234.56', '1 234,56'), ('4200000000', '4.2 миллиарда')):
print(f" score_answer({gt!r}, {pred!r}) = {score_answer(gt, pred, 0.01)}")
Observed
| # |
Input |
Expected |
Actual |
| 1 |
extract_numbers_with_context('4,2') |
4.2 |
42.0 — rescaled 10x, no warning |
| 2 |
score_answer('2602', '2 602') |
1.0 |
0.0 (the 2,602 form scores 1.0) |
| 3 |
detect_unit_in_context('4.2 миллиарда') |
('billion', 1e9) |
(None, 1.0) |
| 4 |
score_answer('4.2', '4,2') |
1.0 |
0.0 |
| 4 |
score_answer('1234.56', '1 234,56') |
1.0 |
0.0 |
| 4 |
score_answer('4200000000', '4.2 миллиарда') |
1.0 |
0.0 |
Five of six semantically equivalent pairs score 0.0.
Root causes
1. reward.py:47 — text_no_commas = text.replace(',', '')
The comma is removed unconditionally before parsing, so it is always a thousands separator and never a decimal separator. In most of Europe, Latin America, and the CIS, 4,2 means 4.2. The result is a 10x error that neither raises nor logs — the worst shape for a scoring bug, because the run completes and the number looks plausible.
2. reward.py:31 — pattern = r'-?\d+\.?\d*%?'
Whitespace-grouped digits (2 602, 1 234,56) are split into separate numbers. Space and non-breaking space are the standard thousands separators in several locales, including the one used by your own officeqa answer style once localised.
3. reward.py:82 — detect_unit_in_context
The scale table is English-only (trillions?, billions?, millions?, thousands? plus bare b/m/k). Non-English scale words return (None, 1.0), so a worded magnitude silently loses its multiplier.
Not locale-specific
For fairness, one failing case is not an i18n issue: number words are unsupported in English too — score_answer('42', 'forty two') is 0.0. Listing it so the scope stays honest; it is a separate feature gap, not part of this report.
Why it matters here
_score_multi_tolerance (src/loop/runner.py:29) is the default scorer and calls score_answer at every tolerance level. Its output feeds update_frontier, so a systematically wrong score does not just misreport accuracy — it changes which programs survive. On a non-English task set the selection pressure is close to noise. This is also the practical blocker for the "Evolution without a benchmark" item in the README: a non-English benchmark cannot be built on a scorer that cannot read non-English numbers.
Open question (why this is a report, not a PR)
The fix needs a decision on how the locale is determined, and there are three reasonable options with different trade-offs:
- Infer per-answer — treat
, as decimal when it is followed by 1-2 digits and there is no other . in the number. Zero configuration, but ambiguous for 1,23.
- Explicit config — a
locale field on the task set. Unambiguous, but adds surface and needs a default.
- Accept both, compare all readings — score as correct if any plausible parse matches. Most permissive, risks false positives.
I have a working implementation of option 1 plus a locale-aware scale table and contract tests of the form "the same value written N ways scores identically". Happy to open a PR against whichever direction you prefer, or to split it so the number-parsing fix lands first and the scale table separately.
Related
Summary
The
multi_tolerancescorer insrc/evaluation/reward.pyassumes English/US number formatting. On non-English answers this does not fail loudly — it silently scores equivalent values as wrong, and in one case silently rescales a number by 10x. Since_score_multi_toleranceis the default scorer and its output drives frontier selection, an evolution run on a non-English task set optimises against a largely meaningless signal.Filing this as a report rather than a PR because the fix involves a design choice (how to decide a locale) that is yours to make — see "Open question" at the end.
Environment
main@36f6f04Reproduction
Save as
repro_locale.pyin the repo root and runpython repro_locale.py:Observed
extract_numbers_with_context('4,2')4.242.0— rescaled 10x, no warningscore_answer('2602', '2 602')1.00.0(the2,602form scores1.0)detect_unit_in_context('4.2 миллиарда')('billion', 1e9)(None, 1.0)score_answer('4.2', '4,2')1.00.0score_answer('1234.56', '1 234,56')1.00.0score_answer('4200000000', '4.2 миллиарда')1.00.0Five of six semantically equivalent pairs score
0.0.Root causes
1.
reward.py:47—text_no_commas = text.replace(',', '')The comma is removed unconditionally before parsing, so it is always a thousands separator and never a decimal separator. In most of Europe, Latin America, and the CIS,
4,2means4.2. The result is a 10x error that neither raises nor logs — the worst shape for a scoring bug, because the run completes and the number looks plausible.2.
reward.py:31—pattern = r'-?\d+\.?\d*%?'Whitespace-grouped digits (
2 602,1 234,56) are split into separate numbers. Space and non-breaking space are the standard thousands separators in several locales, including the one used by your ownofficeqaanswer style once localised.3.
reward.py:82—detect_unit_in_contextThe scale table is English-only (
trillions?,billions?,millions?,thousands?plus bareb/m/k). Non-English scale words return(None, 1.0), so a worded magnitude silently loses its multiplier.Not locale-specific
For fairness, one failing case is not an i18n issue: number words are unsupported in English too —
score_answer('42', 'forty two')is0.0. Listing it so the scope stays honest; it is a separate feature gap, not part of this report.Why it matters here
_score_multi_tolerance(src/loop/runner.py:29) is the default scorer and callsscore_answerat every tolerance level. Its output feedsupdate_frontier, so a systematically wrong score does not just misreport accuracy — it changes which programs survive. On a non-English task set the selection pressure is close to noise. This is also the practical blocker for the "Evolution without a benchmark" item in the README: a non-English benchmark cannot be built on a scorer that cannot read non-English numbers.Open question (why this is a report, not a PR)
The fix needs a decision on how the locale is determined, and there are three reasonable options with different trade-offs:
,as decimal when it is followed by 1-2 digits and there is no other.in the number. Zero configuration, but ambiguous for1,23.localefield on the task set. Unambiguous, but adds surface and needs a default.I have a working implementation of option 1 plus a locale-aware scale table and contract tests of the form "the same value written N ways scores identically". Happy to open a PR against whichever direction you prefer, or to split it so the number-parsing fix lands first and the scale table separately.
Related