collaborators

6 papers

cs.LG2026

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Tao Wang, Shuo Li, Yan Sun +2

Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimi…

cs.CL2026

Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment

Tao Wang, Lipeng Zhu, Jiayong Li +2

Defect grading of power transmission equipment (DGPTE) is crucial to the stability of electric energy transmission. Although existing machine learning methods exhibit strong capabi…

cs.AI2026

Foundations of Top- Decoding For Language Models

Georgy Noarov, Soham Mallick, Tao Wang +5

Top- decoding is a widely used method for sampling from LLMs: at each token, only the largest next-token-probabilities are kept, and the next token is sampled after re-norma…

cs.AI2026

Statistical Early Stopping for Reasoning Models

Yangxinyu Xie, Tao Wang, Soham Mallick +6

While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given…

stat.ML2026

Optimal Decision-Making Based on Prediction Sets

Tao Wang, Edgar Dobriban

Prediction sets can wrap around any ML model to cover unknown test outcomes with a guaranteed probability. Yet, it remains unclear how to use them optimally for downstream decision…

stat.ML2026

Singleton-Optimized Conformal Prediction

Tao Wang, Yan Sun, Edgar Dobriban

Conformal prediction can be used to construct prediction sets that cover the true outcome with a desired probability, but can sometimes lead to large prediction sets that are costl…