Hi — this is unusually careful work. The part that stood out to me is not just the Pokémon framing; it is that the repo treats harness effects as first-class evidence: same tools/memory/prompt/observation format, RAM-checkable predicates, exclusions, and explicit caveats around the 1024-token ceiling.
I pulled a clean checkout and ran the fast validity/score path:
uv run pytest -q tests/test_validity.py tests/test_cli_score.py
.............. [100%]
14 passed in 0.19s
One small adoption question: would you consider adding a short "reproduce one leaderboard row" smoke command to the README or docs?
Something like:
# no ROM / no model key path, just scoring integrity
uv run pytest -q tests/test_validity.py tests/test_cli_score.py
uv run pokebench score --input results.json --scenario s1_exit_pallet --model <one-row>
Or whatever command best fits the current CLI.
Why I ask: the strongest claim here is comparability. A reviewer should be able to go from clean checkout to "I can verify the scoring/schema for one published row" before touching emulator/model runtime. You already have the tests and artifacts; a tiny copy-paste path would make the benchmark much easier to trust and cite.
Hi — this is unusually careful work. The part that stood out to me is not just the Pokémon framing; it is that the repo treats harness effects as first-class evidence: same tools/memory/prompt/observation format, RAM-checkable predicates, exclusions, and explicit caveats around the 1024-token ceiling.
I pulled a clean checkout and ran the fast validity/score path:
One small adoption question: would you consider adding a short "reproduce one leaderboard row" smoke command to the README or docs?
Something like:
Or whatever command best fits the current CLI.
Why I ask: the strongest claim here is comparability. A reviewer should be able to go from clean checkout to "I can verify the scoring/schema for one published row" before touching emulator/model runtime. You already have the tests and artifacts; a tiny copy-paste path would make the benchmark much easier to trust and cite.