5 papers
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Mohsen Hariri, Weicong Chen, Nahal Shahini +11
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algori…
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
Mohsen Hariri, Amirhossein Samandar, Michael Hinczewski +1
Pass is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especially when the number of trials (sampl…
Scorio.jl: A Julia package for ranking stochastic responses
Mohsen Hariri, Michael Hinczewski, Vipin Chaudhary
Scorio.jl is a Julia package for evaluating and ranking systems from repeated responses to shared tasks. It provides a common tensor-based interface for direct score-based, pairwis…
Ranking Reasoning LLMs under Test-Time Scaling
Mohsen Hariri, Michael Hinczewski, Jing Ma +1
Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking models in this regime remains underexplored. We formalize dense benchmark ranking un…
Thermodynamic Performance Limits for Score-Based Diffusion Models
Nathan X. Kodama, Michael Hinczewski
We establish a fundamental connection between score-based diffusion models and non-equilibrium thermodynamics by deriving performance limits based on entropy rates. Our main theore…