6 papers
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
Tao Wang, Shuo Li, Yan Sun +2
Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimi…
Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment
Tao Wang, Lipeng Zhu, Jiayong Li +2
Defect grading of power transmission equipment (DGPTE) is crucial to the stability of electric energy transmission. Although existing machine learning methods exhibit strong capabi…
Foundations of Top- Decoding For Language Models
Georgy Noarov, Soham Mallick, Tao Wang +5
Top- decoding is a widely used method for sampling from LLMs: at each token, only the largest next-token-probabilities are kept, and the next token is sampled after re-norma…
Statistical Early Stopping for Reasoning Models
Yangxinyu Xie, Tao Wang, Soham Mallick +6
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given…
Optimal Decision-Making Based on Prediction Sets
Tao Wang, Edgar Dobriban
Prediction sets can wrap around any ML model to cover unknown test outcomes with a guaranteed probability. Yet, it remains unclear how to use them optimally for downstream decision…
Singleton-Optimized Conformal Prediction
Tao Wang, Yan Sun, Edgar Dobriban
Conformal prediction can be used to construct prediction sets that cover the true outcome with a desired probability, but can sometimes lead to large prediction sets that are costl…