2 papers
cs.RO2026
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
Yanru Wu, Weiduo Yuan, Ang Qi +3
Reinforcement Learning (RL) has shown great potential in refining robotic manipulation policies, yet its efficacy remains strongly bottlenecked by the difficulty of designing gener…
cs.LG2026
FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
Jiajun Xu, Jiageng Mao, Ang Qi +5
Vision Language Models (VLMs) are prone to errors, and identifying where these errors occur is critical for ensuring the reliability and safety of AI systems. In this paper, we pro…