Skip to content
#

inference-time-scaling

Here are 23 public repositories matching this topic...

Evolving agent harnesses: a research program on how far N orchestrated calls of a small model can rival a frontier model. We evolve the harness (structure + prompts) with reflective optimizers + a verified-acceptance gate.

  • Updated Jul 1, 2026
  • Go

Open-source research engineering project for building the end-to-end post-training stack for reasoning language models, including SFT, preference learning, RLHF/RLVR, evaluation, inference-time scaling, and scalable systems for frontier-level reasoning.

  • Updated Jul 18, 2026
  • Jupyter Notebook

How Chain-of-Thought Budgets Induce Overconfidence in LLMs. Investigating Calibration Drift Under Reasoning (CDUR), hypothesis lock-in mechanisms, and dynamic token budgeting via the CABStop optimal stopping algorithm.

  • Updated Jun 16, 2026
  • Python

Official code for "Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration" (SOE, ACL 2026 Oral). A training-free, test-time scaling method that fixes reasoning collapse in LLM mathematical reasoning via orthogonal probing.

  • Updated Jul 7, 2026
  • Python

Improve this page

Add a description, image, and links to the inference-time-scaling topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the inference-time-scaling topic, visit your repo's landing page and select "manage topics."

Learn more