2 papers
cs.CL2025
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
Gonçalo Faria, Noah A. Smith
Increasing test-time computation has emerged as a promising direction for improving language model performance, particularly in scenarios where model finetuning is impractical or i…
cs.AI2025
The Leaderboard Illusion
Shivalika Singh, Yiyang Nan, Alex Wang +10
Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbo…