10 papers
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
Zhepei Wei, Xiao Yang, Kai Sun +12
While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly…
WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
Yuxin Meng, Yuhan Suo, Junjie Wang +9
Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page…
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
Zhepei Wei, Xinyu Zhu, Wei-Lin Chen +3
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the underlying geometry of the res…
Video-Zero: Self-Evolution Video Understanding
Ruixu Zhang, Deyi Ji, Lanyun Zhu +4
Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains…
G-Zero: Self-Play for Open-Ended Generation from Zero Data
Chengsong Huang, Haolin Liu, Tong Zheng +7
Self-evolving LLMs excel in verifiable domains but struggle in open-ended tasks, where reliance on proxy LLM judges introduces capability bottlenecks and reward hacking. To overcom…
A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms
Kejin Yu, Yuhan Sun, Taiqiang Wu +5
The development of high-level autonomous driving (AD) is shifting from perception-centric limitations to a more fundamental bottleneck, namely, a deficit in robust and generalizabl…