2 papers
cs.LG2025
ROC-n-reroll: How verifier imperfection affects test-time scaling
Florian E. Dorner, Yatong Chen, André F. Cruz +2
Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (Bo…
cs.LG2024
Evaluating language models as risk scores
André F. Cruz, Moritz Hardt, Celestine Mendler-Dünner
Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the…