3 papers
cs.RO2025
A Navigation Framework Utilizing Vision-Language Models
Yicheng Duan, Kaiyu tang
Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, un…
cs.CV2025
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Yi-Fan Zhang, Xingyu Lu, Xiao Hu +13
Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on im…
cs.CL2024
Kwai-STaR: Transform LLMs into State-Transition Reasoners
Xingyu Lu, Yuhang Hu, Changyi Liu +12
Mathematical reasoning presents a significant challenge to the cognitive capabilities of LLMs. Various methods have been proposed to enhance the mathematical ability of LLMs. Howev…