1 paper
Qingyuan Wu, Jianheng Liu, Jianye Hao +2
State-of-the-art (SOTA) reinforcement learning (RL) methods have enabled vision-language model (VLM) agents to learn from interaction with online environments without human supervi…