collaborators

5 papers

cs.LG2026

Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training

Peng Cui, Boyao Yang, Jun Zhu

Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), yet prevailing techniques such…

cs.LG2026

Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning

Peng Cui, Boyao Yang, Jun Zhu

Reinforcement learning post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal corr…

cs.AI2026

Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification

Yichi Zhang, Nabeel Seedat, Yinpeng Dong +3

As LLM-powered agents have been used for high-stakes decision-making, such as clinical diagnosis, it becomes critical to develop reliable verification of their decisions to facilit…

cs.CV2024

Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models

Peng Cui, Guande He, Dan Zhang +3

Datasets collected from the open world unavoidably suffer from various forms of randomness or noiseness, leading to the ubiquity of aleatoric (data) uncertainty. Quantifying such u…

cs.AI2024

Accurate and Reliable Predictions with Mutual-Transport Ensemble

Han Liu, Peng Cui, Bingning Wang +2

Deep Neural Networks (DNNs) have achieved remarkable success in a variety of tasks, especially when it comes to prediction accuracy. However, in complex real-world scenarios, parti…