3 papers
cs.CL2025
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
Yingjia Wan, Haochen Tan, Xiao Zhu +9
Evaluating the factuality of long-form generations from Large Language Models (LLMs) remains challenging due to efficiency bottlenecks and reliability concerns. Prior efforts attem…
cs.AI2025
CHARM: Calibrating Reward Models With Chatbot Arena Scores
Xiao Zhu, Chenmien Tan, Pinzhen Chen +4
Reward models (RMs) play a crucial role in Reinforcement Learning from Human Feedback by serving as proxies for human preferences in aligning large language models. However, they s…
cs.CV2025
Regression in EO: Are VLMs Up to the Challenge?
Xizhe Xue, Xiao Xiang Zhu
Earth Observation (EO) data encompass a vast range of remotely sensed information, featuring multi-sensor and multi-temporal, playing an indispensable role in understanding our pla…