Skip to content

Repository files navigation

FORGE: Formally Generated Reasoning Evaluation

An open question about LLM evaluation, built as a tool.

⚠ Current Status: Active Research Phase FORGE is in active research and development. It is not production ready. Current results should be treated as preliminary. We are working toward v1.0 with full scientific validation of all claims. See Known Limitations and Roadmap to v1.0 below.


The Hypothesis

This project starts from a contested scientific question: when a large language model solves a problem correctly, is it doing something that resembles reasoning over internalized structure, or is it pattern-matching against statistical regularities in its training data?

We do not know the answer. Neither does anyone else, with confidence.

FORGE is an instrument designed to produce evidence that bears on this question — not to answer it.

What we hypothesize

We hypothesize that a model's performance profile across difficulty tiers — specifically the relationship between its interpolation score (easy/medium problems) and its extrapolation score (hard/expert problems) — is informative about whether it has internalized mathematical structure or is primarily retrieving from training distribution.

A model that scores well on easy problems and collapses on hard ones is consistent with sophisticated retrieval. A model that degrades gradually across tiers is consistent with something more like generalization. We do not claim this distinction is clean, or that FORGE scores prove either interpretation.

Falsifiability

The hypothesis would be evidence against if models scoring highly on maximum difficulty tiers suffer catastrophic performance degradation when the mathematical structure of a problem is held constant but its surface framing is inverted. That would suggest reliance on surface pattern matching rather than internalized structure. We have not run this experiment systematically. It is a direction for future work.


The Problem with Static Benchmarks

Existing benchmarks have well-documented structural problems:

  • Data contamination: Static question banks are likely present in pre-training corpora, making it difficult to distinguish memorization from reasoning
  • Ceiling effects: Multiple-choice formats are gameable; models exploit linguistic heuristics
  • LLM-as-judge evaluation: Open-ended grading introduces its own hallucination and consistency problems
  • The complexity cliff: Several papers have documented that model accuracy does not degrade gracefully at the boundary of training distribution — it collapses

FORGE addresses these by generating every problem procedurally at runtime from a seed. FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Categories flagged as [RESEARCH] or [FLAGGED] below have not yet achieved the uniqueness threshold required for a full anti-contamination guarantee. Structural similarity to training data — problems that resemble the type of problem seen during training — cannot be ruled out and is a genuine limitation.


The Nerfed Model Problem

AI labs face a quiet incentive problem. A model is released. It scores well on public benchmarks. The scores become part of the marketing. Then, quietly, the model is updated — safety tuning, cost optimization, behavior changes. The published scores no longer reflect what users are actually running.

Fixed benchmarks cannot distinguish between "the model got smarter" and "the model was tuned toward this specific benchmark distribution."

Because every FORGE run is governed by a published seed and every question is generated deterministically at runtime, a score is reproducible. Any researcher can rerun seed 42 six months from now against the same model endpoint and compare. If the score has shifted, the shift is documented.

Whether score changes over model versions reflect capability changes or distribution-specific tuning is something FORGE data may help investigate — though drawing firm conclusions would require controlled conditions we do not currently enforce.

FORGE does not accuse anyone of nerfing. It makes the question more tractable.


How FORGE Works

25 Procedurally Generated Categories

Each category tests a distinct axis of theoretical internal world-models:

# Category Difficulty Range Human Time (Frontier)
1 Arithmetic Chain Composition Kindergarten → Frontier 15 mins
2 Polynomial Root Finding High School → Frontier 20 mins
3 Matrix Determinants Undergraduate → Frontier 45 mins
4 Jordan Normal Form Undergraduate → Frontier 60 mins
5 RLC Circuit Differential Equations High School → Frontier 30 mins
6 Boolean Logic Minimization Undergraduate → Frontier 10 mins
7 Chess Mate-in-N (Endgame) Elementary → Frontier 25 mins
8 Graph Theory: Shortest Path High School → Frontier 15 mins
9 Cryptographic Arithmetic (RSA) Undergraduate → Frontier 40 mins
10 Combinatorics: Stars and Bars High School → Frontier 10 mins
11 Information Theory: Entropy Undergraduate → Frontier 5 mins
12 Signal Processing: DFT Undergraduate → Frontier 20 mins
13 Game Theory: Nim States Elementary → Frontier 5 mins
14 Abstract Algebra: Group Orders Undergraduate → Frontier 15 mins
15 Systems of Linear Equations High School → Frontier 25 mins
16 Modular Exponentiation Middle School → Frontier 10 mins
17 Vector Calculus: Divergence Undergraduate → Frontier 15 mins
18 Geometry: Polygon Properties Middle School → Frontier 15 mins
19 Financial Mathematics High School → Frontier 10 mins
20 Probability: Bayesian Updating Undergraduate → Frontier 15 mins
21 Taylor Series Coefficients Undergraduate → Frontier 25 mins
22 Diophantine Equations High School → Frontier 15 mins
23 Formal Grammars Undergraduate → Frontier 20 mins
24 Quantum State Amplitudes Undergraduate → Frontier 30 mins
25 Algorithmic Trace Execution Middle School → Frontier 20 mins

Deterministic Generation

All problems are generated from a master seed using SHA-256 derived RNGs:

  • Same seed → identical problems across runs, on any machine
  • Different seeds → statistically distinct problems
  • FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Structural similarity to training data is a separate concern and cannot be excluded

Seed Security

FORGE accepts seeds with 2^256 possible input values via SHA-256. The number of genuinely distinct questions per category varies and is documented per category. Question set uniqueness at the 25-category level is substantially larger than any individual category space. Full uniqueness analysis is available at [link to research doc].

Metric Value
Seed input space 2^256 ≈ 1.2 × 10^77
Time to exhaust (1 trillion/sec) ~3.7 × 10^51 years
Atom count in universe ~10^80

Note: The above table describes the seed space. The actual number of distinct questions generated per category is bounded by that category's parameter space, which varies. See Category Status for per-category assessments.

Reproducibility note (v0.2.0): The seed space was increased from 2^64 to 2^256 in this version. Snapshots and results from v0.1.x (which used 2^64 truncation) are not compatible with v0.2.0+ outputs. The previous outputs (5 runs) are not significant enough to preserve. Runtime reproducibility is verified by verify_reproducibility() — if two independent generations with the same seed produce identical output, the result is flagged as reproducible.

Best practices:

  • For public benchmarks: Use a random seed and publish it with your results
  • For contamination testing: Use a seed that is not disclosed in advance
  • For reproducibility: Document the seed alongside your score

If a lab knows your seed, they can pre-compute all problems. Keep seeds private for genuine evaluations.

Computational Grading

FORGE uses deterministic computational verification rather than LLM judges:

  • SymPy: Symbolic equivalence for expressions, equations, and calculus
  • NumPy: Matrix operations and numerical precision
  • python-chess: Game state validation and legal move verification

This grading approach eliminates LLM judge bias but introduces its own edge cases, documented per category. Known grading limitations are flagged in the Category Status table below. The same input always produces the same grade, and the grading logic is inspectable.

Scoring

Per-category score (extrapolation-weighted):

C_c = Σ(α^d · S_{c,d}) / Σ(α^d)

Where:

  • S_{c,d} = accuracy for category c at difficulty d
  • α = 1.5 (extrapolation weight parameter)
  • Higher difficulties contribute exponentially more to the score

FORGE Score: Mean of all per-category scores C_c.

Complexity Cliff Index: The first difficulty tier where accuracy drops more than 30% from the previous tier. This is an observation about where performance degrades — it is not a proof of anything about world models. It is a useful diagnostic number.


Limitations

These are real limitations, not caveats. Read them before drawing conclusions from FORGE scores.

Statistical power by mode:

Mode Questions 95% CI margin Interpretation
nano 5 ±44% Smoke test only. Not meaningful for comparison.
quick 250 ±6.2% Useful for rough ordering. Not sufficient for fine distinctions.
standard 2,500 ±2.0% Adequate for most research comparisons.
full 10,000 ±1.0% High confidence. Expensive.

Quick mode results on cheap models are not meaningful tests of the hypothesis. The statistical power is too low and the models are unlikely to be operating in the regime the hypothesis is about.

Grading edge cases:

  • SymPy: Symbolic simplification can fail on certain trigonometric identities, piecewise functions, and expressions involving special functions. When SymPy cannot simplify a difference to zero, it returns False even if the expressions are mathematically equivalent. This produces false negatives.
  • Numeric tolerance: The 1e-4 relative tolerance used in most numeric categories is a judgment call. Problems with very large or very small answers may grade incorrectly at this threshold.
  • Quantum amplitudes: The answer parser handles ket notation, JSON arrays, and comma-separated complex values. Unusual formatting from models may not parse correctly, producing false negatives.
  • Chess: Puzzle generation uses a minimax search over randomly placed endgame positions. The search is bounded (400 attempts per problem). In rare cases the fallback position is used. The grader accepts any move that forces mate in N, not just the canonical first move.
  • Formal grammars and algorithmic trace: These categories involve string matching and execution trace comparison. Edge cases in whitespace handling and output formatting may affect grading.

What FORGE does not control for:

  • Structural similarity between generated problems and training data. Procedural generation prevents verbatim contamination; it does not prevent a model from having seen many similar problems during training.
  • Prompt sensitivity. The system prompt instructs models to output answers in a specific format. Models that do not follow this format will score lower regardless of whether they computed the correct answer.
  • Reasoning token costs. Models that use extended thinking consume significantly more tokens per problem. Cost estimates in the UI are approximate.

Known Limitations

These are active, specific issues documented with evidence. They are not hypothetical.

Polynomial Roots (category 2) — [FLAGGED]: Problem space of 62.8 million with collision threshold of 7,930. Exhaustive contamination is feasible in hours on consumer hardware. A motivated party could pre-compute all questions for a given seed. Parameter expansion is required before this category can be certified.

RSA Arithmetic (category 9) — [FLAGGED]: Involves large prime operations (modular exponentiation with large moduli) that language models are not designed to perform mentally. This category may test computational token-generation limits rather than reasoning depth, which is inconsistent with the FORGE hypothesis. Under review for redesign or removal.

Boolean Minimization (category 6) — [RESEARCH]: Known grader bug at difficulty 4-5. SymPy's simplify_logic raises TypeError on certain complex Boolean expressions. Fix in progress.

Chess Mate-in-N (category 7) — [RESEARCH]: Generation pool constraints limit true procedural uniqueness at high difficulty levels. The minimax search without alpha-beta pruning is slow for mate-in-4 and mate-in-5. Rewrite in progress.

Algebra Groups (category 14) — [RESEARCH]: Extremely slow generation at high difficulty. Skipped in performance tests. Optimization required before this category can run at scale.

Quantum Amplitudes (category 24) — [RESEARCH]: Answer parser is fragile on non-standard notation. Grading reliability not fully verified across diverse model output formats.

Shannon Entropy (category 11) — [RESEARCH]: Cross-model testing revealed borderline tolerance cases where models producing correct values within rounding error were graded incorrectly. Tolerance thresholds are under review.

Jordan Normal Form (category 4) — [RESEARCH]: At difficulty 5, generated matrices occasionally have degenerate eigenvalue structures that make the problem ambiguous. The grader handles common cases but may miss valid alternative forms.


Roadmap to v1.0

Before FORGE reaches v1.0 with full scientific validation, the following must be completed:

Category Certification Roadmap

Category Current Required for CERTIFIED Status
Polynomial Roots [FLAGGED] Expand parameter space beyond 62.8M unique problems In progress — parameter expansion design
RSA Arithmetic [FLAGGED] Redesign to test reasoning, not computation, or remove Under review
Boolean Minimization [RESEARCH] Fix SymPy simplify_logic TypeError at diff 4-5 Fix in progress
Chess Mate-in-N [RESEARCH] Rewrite with alpha-beta pruning, verify uniqueness at high difficulty Rewrite in progress
Algebra Groups [RESEARCH] Optimize generation speed for high difficulty Not started
Quantum Amplitudes [RESEARCH] Expand answer parser, verify grading across output formats In progress
Shannon Entropy [RESEARCH] Review tolerance thresholds, document edge cases Under review
Jordan Normal Form [RESEARCH] Handle degenerate eigenvalue ambiguity at difficulty 5 Not started
Formal Grammars [RESEARCH] Fix whitespace/formatting edge cases in grading Not started
Algorithmic Trace [RESEARCH] Fix whitespace/formatting edge cases in grading Not started

Infrastructure Roadmap

  • Full uniqueness analysis document with per-category problem space measurements
  • Grader self-consistency verification at 100% for all [RESEARCH] categories
  • Difficulty scaling empirical validation (not just proxy measurement)
  • Cross-seed contamination analysis
  • Publication-grade statistical methodology documentation

Observed Behaviors

Shannon Entropy — log2 hallucination

During evaluation on seed 42, qwen/qwen3-coder-flash produced confident multi-step entropy calculations using incorrect log2 values. Example: log2(107) computed as 6.7175 versus correct 6.7415 (error −0.024). The model used log2(a/b) = log2(a) − log2(b) with systematically wrong log2 values for integers 103, 107, 221, 309, 321, 332, 347. The reasoning trace showed no uncertainty at the point of error. FORGE's deterministic grader flagged all 10 entropy questions as incorrect. Raw response (first failure, truncated): "log2(50/321) = log2(50) - log2(321) = 5.6439 - 8.3219 = -2.6780" (actual: -2.6826).

Shannon Entropy — cross-model non-monotonic failure

On seed 17017656696087371159 (nano mode, 5 questions, Shannon Entropy difficulty 3), 8 frontier models were evaluated. Five passed (DeepSeek V4 Pro, GPT-5.5, Claude Opus 4.7, Nemotron 550B, MiniMax M3). Three failed on the same entropy question (expected 2.7037): Claude Opus 4.8 extracted 2.7038, Claude Opus 4.6 extracted 2.7036 — both within 4th-decimal rounding error of the correct value. Claude Sonnet 4.6 extracted 2.7056, a larger deviation traced to an incorrect intermediate logarithm: ln(0.072650) computed as −2.6439 versus correct −2.6226. Notably, Opus 4.7 passes while 4.6 and 4.8 fail — a non-monotonic pattern across versions. Sonnet 4.6 is the only model with a demonstrably wrong logarithm value; the other two failures are borderline tolerance cases.


Known Issues

These are active limitations in the current implementation.

Chess (category 7): The minimax search for mate-in-4 and mate-in-5 positions can be slow because it searches without alpha-beta pruning. This is a performance issue, not a correctness issue. A future version should add pruning or use a tablebase.

Quantum amplitudes (category 24): The answer parser handles common output formats but will fail on unusual representations. This category has higher false-negative rates than others.

Jordan Normal Form (category 4): At difficulty 5, the generated matrices occasionally have degenerate eigenvalue structures that make the problem ambiguous. The grader handles the most common cases but may miss valid alternative forms.

Statistical independence: Problems within a category at the same difficulty level are generated from independent seeds, but the underlying mathematical structure follows the same parametric distribution. This is not full statistical independence and may affect confidence interval calculations.


Installation

git clone https://github.com/mohamedhossammohamed/FORGE.git
cd FORGE

python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

pip install -r requirements.txt

Quick Start

CLI

# Smoke test — 5 questions, ~30 seconds
python -m forge.cli run \
    --mode nano \
    --model gpt-4o \
    --api-base https://api.openai.com/v1 \
    --api-key sk-your-key-here

# Quick evaluation — 250 questions, ~5 minutes
python -m forge.cli run \
    --mode quick \
    --model gpt-4o \
    --api-base https://api.openai.com/v1 \
    --api-key sk-your-key-here \
    --seed 42

# Standard evaluation — 2,500 questions, ~1 hour
python -m forge.cli run \
    --mode standard \
    --model claude-3-5-sonnet \
    --api-base https://api.anthropic.com/v1 \
    --api-key your-key-here \
    --seed 42

# Run specific categories only
python -m forge.cli run \
    --mode quick \
    --categories arithmetic_chain,matrix_det,game_nim \
    --model gpt-4o \
    --api-base https://api.openai.com/v1 \
    --api-key sk-your-key-here

# List all categories
python -m forge.cli list-categories

Web UI (Researcher Tool)

python -m forge.server
# Opens at http://localhost:7860/researcher.html

The researcher tool provides real-time progress, per-category breakdown, cliff index visualization, multi-model comparison, and snapshot export.


Score Interpretation

Metric What it measures What it does not prove
FORGE Score Weighted accuracy, emphasizing harder problems That the model has an internal world-model
Interpolation Score Accuracy on easy/medium tiers That these problems were in training data
Extrapolation Score Accuracy on hard/expert tiers That these problems were absent from training data
Cliff Index First tier where accuracy drops >30% The cause of the drop

A model with high interpolation and low extrapolation scores is consistent with retrieval-dominant behavior. A model that maintains accuracy across tiers is consistent with generalization. Neither interpretation is proven by the score alone.


Category Status

Every category is assigned a status flag based on current research findings:

  • [FORGE CERTIFIED] — Problem space too large for feasible exhaustive contamination, grader self-consistency verified at 100%, difficulty scaling empirically validated, no known generation edge cases
  • [RESEARCH] — Included but has known limitations: problem space may be feasible to exhaust, grading has known edge cases, or difficulty scaling is proxy-measured not empirically validated
  • [FLAGGED] — Fundamental issue under investigation: problem space is demonstrably exhaustible, grading produces known false positives/negatives, or questions may plausibly exist in training data. Scores excluded from main FORGE score until resolved
# Category Status Notes
1 Arithmetic Chain Composition [FORGE CERTIFIED] Large parameter space, verified grader
2 Polynomial Root Finding [FLAGGED] Problem space of 62.8M with collision threshold of 7,930. Exhaustive contamination feasible in hours on consumer hardware. Parameter expansion required before certification
3 Matrix Determinants [FORGE CERTIFIED] Large continuous parameter space
4 Jordan Normal Form [RESEARCH] Degenerate eigenvalue ambiguity at difficulty 5
5 RLC Circuit Differential Equations [FORGE CERTIFIED] Continuous parameter space
6 Boolean Logic Minimization [RESEARCH] Known grader bug at difficulty 4-5. SymPy simplify_logic TypeError on complex expressions. Fix in progress
7 Chess Mate-in-N [RESEARCH] Generation pool constraints limit true procedural uniqueness at high difficulty. Rewrite in progress
8 Graph Theory: Shortest Path [FORGE CERTIFIED] Large graph parameter space
9 Cryptographic Arithmetic (RSA) [FLAGGED] Involves large prime operations models are not designed to perform mentally. May test computational limits rather than reasoning depth. Inconsistent with FORGE hypothesis. Under review for redesign or removal
10 Combinatorics: Stars and Bars [FORGE CERTIFIED] Large parameter space
11 Information Theory: Entropy [RESEARCH] Borderline tolerance cases observed in cross-model testing. See Observed Behaviors
12 Signal Processing: DFT [FORGE CERTIFIED] Continuous parameter space
13 Game Theory: Nim States [FORGE CERTIFIED] Large state space
14 Abstract Algebra: Group Orders [RESEARCH] Extremely slow generation at high difficulty. Skipped in performance tests. Optimization required
15 Systems of Linear Equations [FORGE CERTIFIED] Large continuous parameter space
16 Modular Exponentiation [FORGE CERTIFIED] Large parameter space
17 Vector Calculus: Divergence [FORGE CERTIFIED] Continuous parameter space
18 Geometry: Polygon Properties [FORGE CERTIFIED] Large continuous parameter space
19 Financial Mathematics [FORGE CERTIFIED] Large parameter space
20 Probability: Bayesian Updating [FORGE CERTIFIED] Large parameter space
21 Taylor Series Coefficients [FORGE CERTIFIED] Large parameter space
22 Diophantine Equations [FORGE CERTIFIED] Large parameter space
23 Formal Grammars [RESEARCH] String matching edge cases in grading
24 Quantum State Amplitudes [RESEARCH] Answer parser fragile on non-standard notation. Grading reliability not fully verified
25 Algorithmic Trace Execution [RESEARCH] Output formatting edge cases in grading

Scores from [FLAGGED] categories are excluded from the main FORGE score until the issue is resolved. Results from [RESEARCH] categories should be interpreted with caution.


Run Modes

Mode Questions Statistical Power Est. Cost Runtime
nano 5 Smoke test only <$0.01 ~30s
quick 250 Low (±6.2% ME) ~$0.10 ~5 mins
standard 2,500 Moderate (±2.0% ME) ~$1.00 ~1 hour
full 10,000 High (±1.0% ME) ~$4.00 ~4 hours

Adding a Category

FORGE is designed for extensibility. To add a new category:

  1. Create forge/categories/your_category.py
  2. Subclass ForgeCategory
  3. Implement the required interface:
from forge.core.generator import ForgeCategory, Problem

class YourCategory(ForgeCategory):

    @property
    def name(self) -> str:
        return "your_category"

    @property
    def display_name(self) -> str:
        return "Your Category Name"

    def get_difficulty_params(self, difficulty: int) -> dict:
        params = {
            1: {"param1": value},
            2: {...},
            # ...
        }
        return params.get(difficulty, params[3])

    def generate(self, difficulty: int, iteration: int) -> Problem:
        rng = self.get_rng(difficulty, iteration)
        params = self.get_difficulty_params(difficulty)
        # Generate problem using rng — do not use any other randomness source
        return Problem(
            question="...",
            answer=...,
            category=self.name,
            difficulty=difficulty,
            iteration=iteration,
            seed=self.base_seed,
            metadata={...},
        )

    def grade(self, prediction: str, answer) -> bool:
        # Use SymPy, NumPy, or exact comparison — not string heuristics
        return prediction == answer
  1. Register in forge/categories/__init__.py

Design principles for categories

  • Procedural generation: Every question must be generated from the RNG, not selected from a bank
  • Computational grading: Use SymPy, NumPy, or domain-specific engines — not LLM judges or string heuristics
  • Parametric difficulty: Difficulty must be controlled by explicit mathematical parameters, not problem selection
  • Document edge cases: If your grader has known failure modes, document them in the category file

Technical Architecture

forge-benchmark/
├── README.md
├── requirements.txt
├── forge/
│   ├── __init__.py
│   ├── core/
│   │   ├── generator.py       # Base class and shared generation interfaces
│   │   ├── grader.py          # Computational grading utilities
│   │   ├── runner.py          # Model API execution, timeout handling, scoring
│   │   └── state.py           # Seed management and deterministic hashing
│   ├── categories/
│   │   ├── __init__.py        # Category registry
│   │   └── [25 category files]
│   ├── cli.py                 # Typer command-line interface
│   ├── server.py              # FastAPI backend with SSE streaming
│   └── ui.py                  # Gradio interface (alternative to server)
└── configs/
    ├── mode_full.json
    ├── mode_standard.json
    └── mode_quick.json

Contributing

Contributions that improve grading precision, add well-designed categories, or fix documented edge cases are welcome.

  1. Fork the repository
  2. Create a feature branch
  3. Implement your change with tests for edge cases
  4. Submit a pull request with a clear description of what the change fixes or adds

Priority areas

  • Alpha-beta pruning for chess puzzle generation (currently unbounded minimax)
  • Additional mathematical domains (topology, number theory, classical mechanics)
  • Improved quantum amplitude parsing for unusual model output formats
  • Formal verification problems (SAT solving, model checking)

Citation

@software{forge2025,
  title={FORGE: Formally Generated Reasoning Evaluation},
  author={Hossam, Mohamed},
  year={2025},
  url={https://github.com/mohamedhossammohamed/FORGE}
}

License

MIT License — see LICENSE for details.


Acknowledgments

Built on SymPy, NumPy, python-chess, and the scientific Python ecosystem. Motivated by Apple ML Research's "Illusion of Thinking" paper and ongoing concerns about benchmark contamination in the field.


Author

Built by Mohamed Hossam, first-year medical student, as an open question about LLM evaluation. Contributions welcome.

@MohamedHz72007


FORGE is an instrument for producing evidence, not asserting conclusions. The question it asks — whether models generalize or interpolate — does not yet have a clean answer. That is why the question is worth asking carefully.

About

FORGE: Formally Generated Reasoning Evaluation — Testing Understanding Over Interpolation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages