4 papers
VLGOR: Visual-Language Knowledge Guided Offline Reinforcement Learning for Generalizable Agents
Pengsen Liu, Maosen Zeng, Nan Tang +4
Combining Large Language Models (LLMs) with Reinforcement Learning (RL) enables agents to interpret language instructions more effectively for task execution. However, LLMs typical…
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
Kaiyuan Li, Jing-Cheng Pang, Yang Yu
Reinforcement learning from verifiable rewards (RLVR) stimulates the thinking processes of large language models (LLMs), substantially enhancing their reasoning abilities on verifi…
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
Jing-Cheng Pang, Kaiyuan Li, Yidi Wang +3
A central challenge in reinforcement learning (RL) is its dependence on extensive real-world interaction data to learn task-specific policies. While recent work demonstrates that l…
WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making
Zhilong Zhang, Ruifeng Chen, Junyin Ye +8
World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate…