An open question about LLM evaluation, built as a tool.
⚠ Current Status: Active Research Phase FORGE is in active research and development. It is not production ready. Current results should be treated as preliminary. We are working toward v1.0 with full scientific validation of all claims. See Known Limitations and Roadmap to v1.0 below.
This project starts from a contested scientific question: when a large language model solves a problem correctly, is it doing something that resembles reasoning over internalized structure, or is it pattern-matching against statistical regularities in its training data?
We do not know the answer. Neither does anyone else, with confidence.
FORGE is an instrument designed to produce evidence that bears on this question — not to answer it.
We hypothesize that a model's performance profile across difficulty tiers — specifically the relationship between its interpolation score (easy/medium problems) and its extrapolation score (hard/expert problems) — is informative about whether it has internalized mathematical structure or is primarily retrieving from training distribution.
A model that scores well on easy problems and collapses on hard ones is consistent with sophisticated retrieval. A model that degrades gradually across tiers is consistent with something more like generalization. We do not claim this distinction is clean, or that FORGE scores prove either interpretation.
The hypothesis would be evidence against if models scoring highly on maximum difficulty tiers suffer catastrophic performance degradation when the mathematical structure of a problem is held constant but its surface framing is inverted. That would suggest reliance on surface pattern matching rather than internalized structure. We have not run this experiment systematically. It is a direction for future work.
Existing benchmarks have well-documented structural problems:
- Data contamination: Static question banks are likely present in pre-training corpora, making it difficult to distinguish memorization from reasoning
- Ceiling effects: Multiple-choice formats are gameable; models exploit linguistic heuristics
- LLM-as-judge evaluation: Open-ended grading introduces its own hallucination and consistency problems
- The complexity cliff: Several papers have documented that model accuracy does not degrade gracefully at the boundary of training distribution — it collapses
FORGE addresses these by generating every problem procedurally at runtime from a seed. FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Categories flagged as [RESEARCH] or [FLAGGED] below have not yet achieved the uniqueness threshold required for a full anti-contamination guarantee. Structural similarity to training data — problems that resemble the type of problem seen during training — cannot be ruled out and is a genuine limitation.
AI labs face a quiet incentive problem. A model is released. It scores well on public benchmarks. The scores become part of the marketing. Then, quietly, the model is updated — safety tuning, cost optimization, behavior changes. The published scores no longer reflect what users are actually running.
Fixed benchmarks cannot distinguish between "the model got smarter" and "the model was tuned toward this specific benchmark distribution."
Because every FORGE run is governed by a published seed and every question is generated deterministically at runtime, a score is reproducible. Any researcher can rerun seed 42 six months from now against the same model endpoint and compare. If the score has shifted, the shift is documented.
Whether score changes over model versions reflect capability changes or distribution-specific tuning is something FORGE data may help investigate — though drawing firm conclusions would require controlled conditions we do not currently enforce.
FORGE does not accuse anyone of nerfing. It makes the question more tractable.
Each category tests a distinct axis of theoretical internal world-models:
| # | Category | Difficulty Range | Human Time (Frontier) |
|---|---|---|---|
| 1 | Arithmetic Chain Composition | Kindergarten → Frontier | 15 mins |
| 2 | Polynomial Root Finding | High School → Frontier | 20 mins |
| 3 | Matrix Determinants | Undergraduate → Frontier | 45 mins |
| 4 | Jordan Normal Form | Undergraduate → Frontier | 60 mins |
| 5 | RLC Circuit Differential Equations | High School → Frontier | 30 mins |
| 6 | Boolean Logic Minimization | Undergraduate → Frontier | 10 mins |
| 7 | Chess Mate-in-N (Endgame) | Elementary → Frontier | 25 mins |
| 8 | Graph Theory: Shortest Path | High School → Frontier | 15 mins |
| 9 | Cryptographic Arithmetic (RSA) | Undergraduate → Frontier | 40 mins |
| 10 | Combinatorics: Stars and Bars | High School → Frontier | 10 mins |
| 11 | Information Theory: Entropy | Undergraduate → Frontier | 5 mins |
| 12 | Signal Processing: DFT | Undergraduate → Frontier | 20 mins |
| 13 | Game Theory: Nim States | Elementary → Frontier | 5 mins |
| 14 | Abstract Algebra: Group Orders | Undergraduate → Frontier | 15 mins |
| 15 | Systems of Linear Equations | High School → Frontier | 25 mins |
| 16 | Modular Exponentiation | Middle School → Frontier | 10 mins |
| 17 | Vector Calculus: Divergence | Undergraduate → Frontier | 15 mins |
| 18 | Geometry: Polygon Properties | Middle School → Frontier | 15 mins |
| 19 | Financial Mathematics | High School → Frontier | 10 mins |
| 20 | Probability: Bayesian Updating | Undergraduate → Frontier | 15 mins |
| 21 | Taylor Series Coefficients | Undergraduate → Frontier | 25 mins |
| 22 | Diophantine Equations | High School → Frontier | 15 mins |
| 23 | Formal Grammars | Undergraduate → Frontier | 20 mins |
| 24 | Quantum State Amplitudes | Undergraduate → Frontier | 30 mins |
| 25 | Algorithmic Trace Execution | Middle School → Frontier | 20 mins |
All problems are generated from a master seed using SHA-256 derived RNGs:
- Same seed → identical problems across runs, on any machine
- Different seeds → statistically distinct problems
- FORGE uses procedural generation to make verbatim contamination statistically unlikely at normal evaluation scales. Individual category problem spaces vary significantly in size. Structural similarity to training data is a separate concern and cannot be excluded
FORGE accepts seeds with 2^256 possible input values via SHA-256. The number of genuinely distinct questions per category varies and is documented per category. Question set uniqueness at the 25-category level is substantially larger than any individual category space. Full uniqueness analysis is available at [link to research doc].
| Metric | Value |
|---|---|
| Seed input space | 2^256 ≈ 1.2 × 10^77 |
| Time to exhaust (1 trillion/sec) | ~3.7 × 10^51 years |
| Atom count in universe | ~10^80 |
Note: The above table describes the seed space. The actual number of distinct questions generated per category is bounded by that category's parameter space, which varies. See Category Status for per-category assessments.
Reproducibility note (v0.2.0): The seed space was increased from 2^64 to 2^256
in this version. Snapshots and results from v0.1.x (which used 2^64 truncation) are
not compatible with v0.2.0+ outputs. The previous outputs (5 runs) are not significant
enough to preserve. Runtime reproducibility is verified by verify_reproducibility() —
if two independent generations with the same seed produce identical output, the result
is flagged as reproducible.
Best practices:
- For public benchmarks: Use a random seed and publish it with your results
- For contamination testing: Use a seed that is not disclosed in advance
- For reproducibility: Document the seed alongside your score
If a lab knows your seed, they can pre-compute all problems. Keep seeds private for genuine evaluations.
FORGE uses deterministic computational verification rather than LLM judges:
- SymPy: Symbolic equivalence for expressions, equations, and calculus
- NumPy: Matrix operations and numerical precision
- python-chess: Game state validation and legal move verification
This grading approach eliminates LLM judge bias but introduces its own edge cases, documented per category. Known grading limitations are flagged in the Category Status table below. The same input always produces the same grade, and the grading logic is inspectable.
Per-category score (extrapolation-weighted):
C_c = Σ(α^d · S_{c,d}) / Σ(α^d)
Where:
S_{c,d}= accuracy for categorycat difficultydα = 1.5(extrapolation weight parameter)- Higher difficulties contribute exponentially more to the score
FORGE Score: Mean of all per-category scores C_c.
Complexity Cliff Index: The first difficulty tier where accuracy drops more than 30% from the previous tier. This is an observation about where performance degrades — it is not a proof of anything about world models. It is a useful diagnostic number.
These are real limitations, not caveats. Read them before drawing conclusions from FORGE scores.
Statistical power by mode:
| Mode | Questions | 95% CI margin | Interpretation |
|---|---|---|---|
nano |
5 | ±44% | Smoke test only. Not meaningful for comparison. |
quick |
250 | ±6.2% | Useful for rough ordering. Not sufficient for fine distinctions. |
standard |
2,500 | ±2.0% | Adequate for most research comparisons. |
full |
10,000 | ±1.0% | High confidence. Expensive. |
Quick mode results on cheap models are not meaningful tests of the hypothesis. The statistical power is too low and the models are unlikely to be operating in the regime the hypothesis is about.
Grading edge cases:
- SymPy: Symbolic simplification can fail on certain trigonometric identities, piecewise functions, and expressions involving special functions. When SymPy cannot simplify a difference to zero, it returns False even if the expressions are mathematically equivalent. This produces false negatives.
- Numeric tolerance: The 1e-4 relative tolerance used in most numeric categories is a judgment call. Problems with very large or very small answers may grade incorrectly at this threshold.
- Quantum amplitudes: The answer parser handles ket notation, JSON arrays, and comma-separated complex values. Unusual formatting from models may not parse correctly, producing false negatives.
- Chess: Puzzle generation uses a minimax search over randomly placed endgame positions. The search is bounded (400 attempts per problem). In rare cases the fallback position is used. The grader accepts any move that forces mate in N, not just the canonical first move.
- Formal grammars and algorithmic trace: These categories involve string matching and execution trace comparison. Edge cases in whitespace handling and output formatting may affect grading.
What FORGE does not control for:
- Structural similarity between generated problems and training data. Procedural generation prevents verbatim contamination; it does not prevent a model from having seen many similar problems during training.
- Prompt sensitivity. The system prompt instructs models to output answers in a specific format. Models that do not follow this format will score lower regardless of whether they computed the correct answer.
- Reasoning token costs. Models that use extended thinking consume significantly more tokens per problem. Cost estimates in the UI are approximate.
These are active, specific issues documented with evidence. They are not hypothetical.
Polynomial Roots (category 2) — [FLAGGED]: Problem space of 62.8 million with collision threshold of 7,930. Exhaustive contamination is feasible in hours on consumer hardware. A motivated party could pre-compute all questions for a given seed. Parameter expansion is required before this category can be certified.
RSA Arithmetic (category 9) — [FLAGGED]: Involves large prime operations (modular exponentiation with large moduli) that language models are not designed to perform mentally. This category may test computational token-generation limits rather than reasoning depth, which is inconsistent with the FORGE hypothesis. Under review for redesign or removal.
Boolean Minimization (category 6) — [RESEARCH]: Known grader bug at difficulty 4-5. SymPy's simplify_logic raises TypeError on certain complex Boolean expressions. Fix in progress.
Chess Mate-in-N (category 7) — [RESEARCH]: Generation pool constraints limit true procedural uniqueness at high difficulty levels. The minimax search without alpha-beta pruning is slow for mate-in-4 and mate-in-5. Rewrite in progress.
Algebra Groups (category 14) — [RESEARCH]: Extremely slow generation at high difficulty. Skipped in performance tests. Optimization required before this category can run at scale.
Quantum Amplitudes (category 24) — [RESEARCH]: Answer parser is fragile on non-standard notation. Grading reliability not fully verified across diverse model output formats.
Shannon Entropy (category 11) — [RESEARCH]: Cross-model testing revealed borderline tolerance cases where models producing correct values within rounding error were graded incorrectly. Tolerance thresholds are under review.
Jordan Normal Form (category 4) — [RESEARCH]: At difficulty 5, generated matrices occasionally have degenerate eigenvalue structures that make the problem ambiguous. The grader handles common cases but may miss valid alternative forms.
Before FORGE reaches v1.0 with full scientific validation, the following must be completed:
| Category | Current | Required for CERTIFIED | Status |
|---|---|---|---|
| Polynomial Roots | [FLAGGED] | Expand parameter space beyond 62.8M unique problems | In progress — parameter expansion design |
| RSA Arithmetic | [FLAGGED] | Redesign to test reasoning, not computation, or remove | Under review |
| Boolean Minimization | [RESEARCH] | Fix SymPy simplify_logic TypeError at diff 4-5 | Fix in progress |
| Chess Mate-in-N | [RESEARCH] | Rewrite with alpha-beta pruning, verify uniqueness at high difficulty | Rewrite in progress |
| Algebra Groups | [RESEARCH] | Optimize generation speed for high difficulty | Not started |
| Quantum Amplitudes | [RESEARCH] | Expand answer parser, verify grading across output formats | In progress |
| Shannon Entropy | [RESEARCH] | Review tolerance thresholds, document edge cases | Under review |
| Jordan Normal Form | [RESEARCH] | Handle degenerate eigenvalue ambiguity at difficulty 5 | Not started |
| Formal Grammars | [RESEARCH] | Fix whitespace/formatting edge cases in grading | Not started |
| Algorithmic Trace | [RESEARCH] | Fix whitespace/formatting edge cases in grading | Not started |
- Full uniqueness analysis document with per-category problem space measurements
- Grader self-consistency verification at 100% for all [RESEARCH] categories
- Difficulty scaling empirical validation (not just proxy measurement)
- Cross-seed contamination analysis
- Publication-grade statistical methodology documentation
Shannon Entropy — log2 hallucination
During evaluation on seed 42, qwen/qwen3-coder-flash produced confident multi-step entropy calculations using incorrect log2 values. Example: log2(107) computed as 6.7175 versus correct 6.7415 (error −0.024). The model used log2(a/b) = log2(a) − log2(b) with systematically wrong log2 values for integers 103, 107, 221, 309, 321, 332, 347. The reasoning trace showed no uncertainty at the point of error. FORGE's deterministic grader flagged all 10 entropy questions as incorrect. Raw response (first failure, truncated): "log2(50/321) = log2(50) - log2(321) = 5.6439 - 8.3219 = -2.6780" (actual: -2.6826).
Shannon Entropy — cross-model non-monotonic failure
On seed 17017656696087371159 (nano mode, 5 questions, Shannon Entropy difficulty 3), 8 frontier models were evaluated. Five passed (DeepSeek V4 Pro, GPT-5.5, Claude Opus 4.7, Nemotron 550B, MiniMax M3). Three failed on the same entropy question (expected 2.7037): Claude Opus 4.8 extracted 2.7038, Claude Opus 4.6 extracted 2.7036 — both within 4th-decimal rounding error of the correct value. Claude Sonnet 4.6 extracted 2.7056, a larger deviation traced to an incorrect intermediate logarithm: ln(0.072650) computed as −2.6439 versus correct −2.6226. Notably, Opus 4.7 passes while 4.6 and 4.8 fail — a non-monotonic pattern across versions. Sonnet 4.6 is the only model with a demonstrably wrong logarithm value; the other two failures are borderline tolerance cases.
These are active limitations in the current implementation.
Chess (category 7): The minimax search for mate-in-4 and mate-in-5 positions can be slow because it searches without alpha-beta pruning. This is a performance issue, not a correctness issue. A future version should add pruning or use a tablebase.
Quantum amplitudes (category 24): The answer parser handles common output formats but will fail on unusual representations. This category has higher false-negative rates than others.
Jordan Normal Form (category 4): At difficulty 5, the generated matrices occasionally have degenerate eigenvalue structures that make the problem ambiguous. The grader handles the most common cases but may miss valid alternative forms.
Statistical independence: Problems within a category at the same difficulty level are generated from independent seeds, but the underlying mathematical structure follows the same parametric distribution. This is not full statistical independence and may affect confidence interval calculations.
git clone https://github.com/mohamedhossammohamed/FORGE.git
cd FORGE
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt# Smoke test — 5 questions, ~30 seconds
python -m forge.cli run \
--mode nano \
--model gpt-4o \
--api-base https://api.openai.com/v1 \
--api-key sk-your-key-here
# Quick evaluation — 250 questions, ~5 minutes
python -m forge.cli run \
--mode quick \
--model gpt-4o \
--api-base https://api.openai.com/v1 \
--api-key sk-your-key-here \
--seed 42
# Standard evaluation — 2,500 questions, ~1 hour
python -m forge.cli run \
--mode standard \
--model claude-3-5-sonnet \
--api-base https://api.anthropic.com/v1 \
--api-key your-key-here \
--seed 42
# Run specific categories only
python -m forge.cli run \
--mode quick \
--categories arithmetic_chain,matrix_det,game_nim \
--model gpt-4o \
--api-base https://api.openai.com/v1 \
--api-key sk-your-key-here
# List all categories
python -m forge.cli list-categoriespython -m forge.server
# Opens at http://localhost:7860/researcher.htmlThe researcher tool provides real-time progress, per-category breakdown, cliff index visualization, multi-model comparison, and snapshot export.
| Metric | What it measures | What it does not prove |
|---|---|---|
| FORGE Score | Weighted accuracy, emphasizing harder problems | That the model has an internal world-model |
| Interpolation Score | Accuracy on easy/medium tiers | That these problems were in training data |
| Extrapolation Score | Accuracy on hard/expert tiers | That these problems were absent from training data |
| Cliff Index | First tier where accuracy drops >30% | The cause of the drop |
A model with high interpolation and low extrapolation scores is consistent with retrieval-dominant behavior. A model that maintains accuracy across tiers is consistent with generalization. Neither interpretation is proven by the score alone.
Every category is assigned a status flag based on current research findings:
- [FORGE CERTIFIED] — Problem space too large for feasible exhaustive contamination, grader self-consistency verified at 100%, difficulty scaling empirically validated, no known generation edge cases
- [RESEARCH] — Included but has known limitations: problem space may be feasible to exhaust, grading has known edge cases, or difficulty scaling is proxy-measured not empirically validated
- [FLAGGED] — Fundamental issue under investigation: problem space is demonstrably exhaustible, grading produces known false positives/negatives, or questions may plausibly exist in training data. Scores excluded from main FORGE score until resolved
| # | Category | Status | Notes |
|---|---|---|---|
| 1 | Arithmetic Chain Composition | [FORGE CERTIFIED] | Large parameter space, verified grader |
| 2 | Polynomial Root Finding | [FLAGGED] | Problem space of 62.8M with collision threshold of 7,930. Exhaustive contamination feasible in hours on consumer hardware. Parameter expansion required before certification |
| 3 | Matrix Determinants | [FORGE CERTIFIED] | Large continuous parameter space |
| 4 | Jordan Normal Form | [RESEARCH] | Degenerate eigenvalue ambiguity at difficulty 5 |
| 5 | RLC Circuit Differential Equations | [FORGE CERTIFIED] | Continuous parameter space |
| 6 | Boolean Logic Minimization | [RESEARCH] | Known grader bug at difficulty 4-5. SymPy simplify_logic TypeError on complex expressions. Fix in progress |
| 7 | Chess Mate-in-N | [RESEARCH] | Generation pool constraints limit true procedural uniqueness at high difficulty. Rewrite in progress |
| 8 | Graph Theory: Shortest Path | [FORGE CERTIFIED] | Large graph parameter space |
| 9 | Cryptographic Arithmetic (RSA) | [FLAGGED] | Involves large prime operations models are not designed to perform mentally. May test computational limits rather than reasoning depth. Inconsistent with FORGE hypothesis. Under review for redesign or removal |
| 10 | Combinatorics: Stars and Bars | [FORGE CERTIFIED] | Large parameter space |
| 11 | Information Theory: Entropy | [RESEARCH] | Borderline tolerance cases observed in cross-model testing. See Observed Behaviors |
| 12 | Signal Processing: DFT | [FORGE CERTIFIED] | Continuous parameter space |
| 13 | Game Theory: Nim States | [FORGE CERTIFIED] | Large state space |
| 14 | Abstract Algebra: Group Orders | [RESEARCH] | Extremely slow generation at high difficulty. Skipped in performance tests. Optimization required |
| 15 | Systems of Linear Equations | [FORGE CERTIFIED] | Large continuous parameter space |
| 16 | Modular Exponentiation | [FORGE CERTIFIED] | Large parameter space |
| 17 | Vector Calculus: Divergence | [FORGE CERTIFIED] | Continuous parameter space |
| 18 | Geometry: Polygon Properties | [FORGE CERTIFIED] | Large continuous parameter space |
| 19 | Financial Mathematics | [FORGE CERTIFIED] | Large parameter space |
| 20 | Probability: Bayesian Updating | [FORGE CERTIFIED] | Large parameter space |
| 21 | Taylor Series Coefficients | [FORGE CERTIFIED] | Large parameter space |
| 22 | Diophantine Equations | [FORGE CERTIFIED] | Large parameter space |
| 23 | Formal Grammars | [RESEARCH] | String matching edge cases in grading |
| 24 | Quantum State Amplitudes | [RESEARCH] | Answer parser fragile on non-standard notation. Grading reliability not fully verified |
| 25 | Algorithmic Trace Execution | [RESEARCH] | Output formatting edge cases in grading |
Scores from [FLAGGED] categories are excluded from the main FORGE score until the issue is resolved. Results from [RESEARCH] categories should be interpreted with caution.
| Mode | Questions | Statistical Power | Est. Cost | Runtime |
|---|---|---|---|---|
nano |
5 | Smoke test only | <$0.01 | ~30s |
quick |
250 | Low (±6.2% ME) | ~$0.10 | ~5 mins |
standard |
2,500 | Moderate (±2.0% ME) | ~$1.00 | ~1 hour |
full |
10,000 | High (±1.0% ME) | ~$4.00 | ~4 hours |
FORGE is designed for extensibility. To add a new category:
- Create
forge/categories/your_category.py - Subclass
ForgeCategory - Implement the required interface:
from forge.core.generator import ForgeCategory, Problem
class YourCategory(ForgeCategory):
@property
def name(self) -> str:
return "your_category"
@property
def display_name(self) -> str:
return "Your Category Name"
def get_difficulty_params(self, difficulty: int) -> dict:
params = {
1: {"param1": value},
2: {...},
# ...
}
return params.get(difficulty, params[3])
def generate(self, difficulty: int, iteration: int) -> Problem:
rng = self.get_rng(difficulty, iteration)
params = self.get_difficulty_params(difficulty)
# Generate problem using rng — do not use any other randomness source
return Problem(
question="...",
answer=...,
category=self.name,
difficulty=difficulty,
iteration=iteration,
seed=self.base_seed,
metadata={...},
)
def grade(self, prediction: str, answer) -> bool:
# Use SymPy, NumPy, or exact comparison — not string heuristics
return prediction == answer- Register in
forge/categories/__init__.py
- Procedural generation: Every question must be generated from the RNG, not selected from a bank
- Computational grading: Use SymPy, NumPy, or domain-specific engines — not LLM judges or string heuristics
- Parametric difficulty: Difficulty must be controlled by explicit mathematical parameters, not problem selection
- Document edge cases: If your grader has known failure modes, document them in the category file
forge-benchmark/
├── README.md
├── requirements.txt
├── forge/
│ ├── __init__.py
│ ├── core/
│ │ ├── generator.py # Base class and shared generation interfaces
│ │ ├── grader.py # Computational grading utilities
│ │ ├── runner.py # Model API execution, timeout handling, scoring
│ │ └── state.py # Seed management and deterministic hashing
│ ├── categories/
│ │ ├── __init__.py # Category registry
│ │ └── [25 category files]
│ ├── cli.py # Typer command-line interface
│ ├── server.py # FastAPI backend with SSE streaming
│ └── ui.py # Gradio interface (alternative to server)
└── configs/
├── mode_full.json
├── mode_standard.json
└── mode_quick.json
Contributions that improve grading precision, add well-designed categories, or fix documented edge cases are welcome.
- Fork the repository
- Create a feature branch
- Implement your change with tests for edge cases
- Submit a pull request with a clear description of what the change fixes or adds
- Alpha-beta pruning for chess puzzle generation (currently unbounded minimax)
- Additional mathematical domains (topology, number theory, classical mechanics)
- Improved quantum amplitude parsing for unusual model output formats
- Formal verification problems (SAT solving, model checking)
@software{forge2025,
title={FORGE: Formally Generated Reasoning Evaluation},
author={Hossam, Mohamed},
year={2025},
url={https://github.com/mohamedhossammohamed/FORGE}
}MIT License — see LICENSE for details.
Built on SymPy, NumPy, python-chess, and the scientific Python ecosystem. Motivated by Apple ML Research's "Illusion of Thinking" paper and ongoing concerns about benchmark contamination in the field.
Built by Mohamed Hossam, first-year medical student, as an open question about LLM evaluation. Contributions welcome.
FORGE is an instrument for producing evidence, not asserting conclusions. The question it asks — whether models generalize or interpolate — does not yet have a clean answer. That is why the question is worth asking carefully.