This repo implements the full pipeline described in the proposal:
- sample prompts from benchmark datasets
- run an open-source LLM token-by-token while collecting per-token signals
- fit a Gaussian-emission HMM to the feature time series
- build a "reasoning map" (latent states + transitions)
- evaluate early forecasting + detour states
python -m venv .venv
source .venv/bin/activate # (Windows: .venv\Scripts\activate)
pip install -r requirements.txtpython scripts/run_tracing.py \
--model mistralai/Mistral-7B-Instruct-v0.2 \
--task gsm8k \
--split test \
--n 200 \
--max-new-tokens 128 \
--temperature 0.7 \
--top-p 0.95 \
--seed 0 \
--out runs/mistral_gsm8k_t0p7This creates runs/.../traces/*.npz plus a manifest.jsonl.
python scripts/train_hmm.py \
--run-dir runs/mistral_gsm8k_t0p7 \
--k-min 3 --k-max 12 \
--cov full \
--n-init 5 \
--max-iter 200Outputs:
runs/.../hmm/model.pklruns/.../hmm/bic.jsonruns/.../hmm/states/*.npz(Viterbi + posteriors)
python scripts/evaluate_forecasting.py \
--run-dir runs/mistral_gsm8k_t0p7 \
--ks 5 10 20 40python scripts/evaluate_detours.py --run-dir runs/mistral_gsm8k_t0p7-
The tracing loop is custom (not
generate) so we can reliably collect: entropy, chosen-token logprob, hidden-state drift/cosine/norm/variance, and attention entropy/max (last layer, averaged over heads). -
"Hallucination" is implemented as a configurable placeholder: by default, it’s
not is_correctfor QA-style tasks. You can plug in retrieval-verified claim checking later viaOutcomeEvaluator.
MIT