Skip to content

Question: tiny reproduce-one-row smoke path for leaderboard trust? #1

Description

@alkhunizan

Hi — this is unusually careful work. The part that stood out to me is not just the Pokémon framing; it is that the repo treats harness effects as first-class evidence: same tools/memory/prompt/observation format, RAM-checkable predicates, exclusions, and explicit caveats around the 1024-token ceiling.

I pulled a clean checkout and ran the fast validity/score path:

uv run pytest -q tests/test_validity.py tests/test_cli_score.py
.............. [100%]
14 passed in 0.19s

One small adoption question: would you consider adding a short "reproduce one leaderboard row" smoke command to the README or docs?

Something like:

# no ROM / no model key path, just scoring integrity
uv run pytest -q tests/test_validity.py tests/test_cli_score.py
uv run pokebench score --input results.json --scenario s1_exit_pallet --model <one-row>

Or whatever command best fits the current CLI.

Why I ask: the strongest claim here is comparability. A reviewer should be able to go from clean checkout to "I can verify the scoring/schema for one published row" before touching emulator/model runtime. You already have the tests and artifacts; a tiny copy-paste path would make the benchmark much easier to trust and cite.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions