3 papers
cs.LG2026
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Mohsen Hariri, Weicong Chen, Nahal Shahini +11
Large language models can solve harder reasoning problems with more inference-time compute. The term "test-time scaling," however, covers several inference algorithms: extending de…
cs.MS2026
Scorio.jl: A Julia package for ranking stochastic responses
Mohsen Hariri, Michael Hinczewski, Vipin Chaudhary
Scorio.jl is a Julia package for evaluating and ranking systems from repeated responses to shared tasks. It provides a common tensor-based interface for direct score-based, pairwis…
cs.LG2025
Thermodynamic Performance Limits for Score-Based Diffusion Models
Nathan X. Kodama, Michael Hinczewski
We establish a fundamental connection between score-based diffusion models and non-equilibrium thermodynamics by deriving performance limits based on entropy rates. Our main theore…