5 papers
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
Peng Cui, Boyao Yang, Jun Zhu
Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), yet prevailing techniques such…
Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning
Peng Cui, Boyao Yang, Jun Zhu
Reinforcement learning post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal corr…
Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification
Yichi Zhang, Nabeel Seedat, Yinpeng Dong +3
As LLM-powered agents have been used for high-stakes decision-making, such as clinical diagnosis, it becomes critical to develop reliable verification of their decisions to facilit…
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
Peng Cui, Guande He, Dan Zhang +3
Datasets collected from the open world unavoidably suffer from various forms of randomness or noiseness, leading to the ubiquity of aleatoric (data) uncertainty. Quantifying such u…
Accurate and Reliable Predictions with Mutual-Transport Ensemble
Han Liu, Peng Cui, Bingning Wang +2
Deep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, especially when it comes to prediction accuracy. However, in complex real-world scenarios, parti…