JAMO — OCR-mediated Hangul rendering benchmark. Code companion to huggingface.co/datasets/Nasser4963/jamo-gold
-
Updated
Aug 17, 2026 - Python
JAMO — OCR-mediated Hangul rendering benchmark. Code companion to huggingface.co/datasets/Nasser4963/jamo-gold
Code, data, and raw model outputs for "Failure Modes in Perturbation-Based Measurement of Language Model Reliability". evaluation/reproduce.py re-derives every reported number offline, no API key required.
A controlled benchmark of eight React state management configurations, with the full replication package: cross-hardware, cross-engine, and a row sweep from 10 to 1000 rows.
Measurement validity, construct validity, and unsupported claims derived from AI agent telemetry.
To associate your repository with the measurement-validity topic, visit your repo's landing page and select "manage topics."