4 papers
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
Tao Wang, Shuo Li, Yan Sun +2
Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimi…
Foundations of Top- Decoding For Language Models
Georgy Noarov, Soham Mallick, Tao Wang +5
Top- decoding is a widely used method for sampling from LLMs: at each token, only the largest next-token-probabilities are kept, and the next token is sampled after re-norma…
Statistical Early Stopping for Reasoning Models
Yangxinyu Xie, Tao Wang, Soham Mallick +6
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given…
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
Sunay Joshi, Yan Sun, Hamed Hassani +1
As generative AI systems are increasingly deployed in real-world applications, regulating multiple dimensions of model behavior has become essential. We focus on test-time filterin…