2 papers
cs.AI2026
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Xuqing Yang, Yi Yuan, Shanzhe Lei +1
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing R…
cs.CV2025
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Lei Li, Yuancheng Wei, Zhihui Xie +9
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current…