Skip to content

Scorer silently mis-scores non-English number formats (decimal comma rescales 10x) #55

Description

@stalker023reg

Summary

The multi_tolerance scorer in src/evaluation/reward.py assumes English/US number formatting. On non-English answers this does not fail loudly — it silently scores equivalent values as wrong, and in one case silently rescales a number by 10x. Since _score_multi_tolerance is the default scorer and its output drives frontier selection, an evolution run on a non-English task set optimises against a largely meaningless signal.

Filing this as a report rather than a PR because the fix involves a design choice (how to decide a locale) that is yours to make — see "Open question" at the end.

Environment

  • main @ 36f6f04
  • Python 3.12, Linux
  • No configuration; default scorer path

Reproduction

Save as repro_locale.py in the repo root and run python repro_locale.py:

from src.evaluation.reward import (
    score_answer,
    extract_numbers_with_context,
    detect_unit_in_context,
)

print("1) decimal comma is silently rescaled")
print("   extract_numbers_with_context('4,2') ->", extract_numbers_with_context('4,2'))

print("2) space thousands separator")
print("   score_answer('2602', '2 602') ->", score_answer('2602', '2 602'))
print("   score_answer('2602', '2,602') ->", score_answer('2602', '2,602'))

print("3) non-English scale words")
for c in ('4.2 billion', '4.2 миллиарда', '543 миллиона'):
    print(f"   detect_unit_in_context({c!r}) -> {detect_unit_in_context(c)}")

print("4) end-to-end")
for gt, pred in (('4.2', '4,2'), ('1234.56', '1 234,56'), ('4200000000', '4.2 миллиарда')):
    print(f"   score_answer({gt!r}, {pred!r}) = {score_answer(gt, pred, 0.01)}")

Observed

# Input Expected Actual
1 extract_numbers_with_context('4,2') 4.2 42.0 — rescaled 10x, no warning
2 score_answer('2602', '2 602') 1.0 0.0 (the 2,602 form scores 1.0)
3 detect_unit_in_context('4.2 миллиарда') ('billion', 1e9) (None, 1.0)
4 score_answer('4.2', '4,2') 1.0 0.0
4 score_answer('1234.56', '1 234,56') 1.0 0.0
4 score_answer('4200000000', '4.2 миллиарда') 1.0 0.0

Five of six semantically equivalent pairs score 0.0.

Root causes

1. reward.py:47 — text_no_commas = text.replace(',', '')

The comma is removed unconditionally before parsing, so it is always a thousands separator and never a decimal separator. In most of Europe, Latin America, and the CIS, 4,2 means 4.2. The result is a 10x error that neither raises nor logs — the worst shape for a scoring bug, because the run completes and the number looks plausible.

2. reward.py:31 — pattern = r'-?\d+\.?\d*%?'

Whitespace-grouped digits (2 602, 1 234,56) are split into separate numbers. Space and non-breaking space are the standard thousands separators in several locales, including the one used by your own officeqa answer style once localised.

3. reward.py:82 — detect_unit_in_context

The scale table is English-only (trillions?, billions?, millions?, thousands? plus bare b/m/k). Non-English scale words return (None, 1.0), so a worded magnitude silently loses its multiplier.

Not locale-specific

For fairness, one failing case is not an i18n issue: number words are unsupported in English too — score_answer('42', 'forty two') is 0.0. Listing it so the scope stays honest; it is a separate feature gap, not part of this report.

Why it matters here

_score_multi_tolerance (src/loop/runner.py:29) is the default scorer and calls score_answer at every tolerance level. Its output feeds update_frontier, so a systematically wrong score does not just misreport accuracy — it changes which programs survive. On a non-English task set the selection pressure is close to noise. This is also the practical blocker for the "Evolution without a benchmark" item in the README: a non-English benchmark cannot be built on a scorer that cannot read non-English numbers.

Open question (why this is a report, not a PR)

The fix needs a decision on how the locale is determined, and there are three reasonable options with different trade-offs:

  1. Infer per-answer — treat , as decimal when it is followed by 1-2 digits and there is no other . in the number. Zero configuration, but ambiguous for 1,23.
  2. Explicit config — a locale field on the task set. Unambiguous, but adds surface and needs a default.
  3. Accept both, compare all readings — score as correct if any plausible parse matches. Most permissive, risks false positives.

I have a working implementation of option 1 plus a locale-aware scale table and contract tests of the form "the same value written N ways scores identically". Happy to open a PR against whichever direction you prefer, or to split it so the number-parsing fix lands first and the scale table separately.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions