5 papers
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
Sai Wang, Yu Wu, Zhongwen Xu
The pursuit of artificial agents that can learn to master complex environments has led to remarkable successes, yet prevailing deep reinforcement learning methods often rely on imm…
Single-stream Policy Optimization
Zhongwen Xu, Zihan Ding
We revisit policy-gradient optimization for Large Language Models (LLMs) from a single-stream perspective. Prevailing group-based methods like GRPO reduce variance with on-the-fly…
Understanding Tool-Integrated Reasoning
Heng Lin, Zhongwen Xu
We study why Tool-Integrated Reasoning (TIR) makes Large Language Models (LLMs) more capable. While LLMs integrated with tools like Python code interpreters show great promise, a p…
Agents Play Thousands of 3D Video Games
Zhongwen Xu, Xianliang Wang, Siyi Li +4
We present PORTAL, a novel framework for developing artificial intelligence agents capable of playing thousands of 3D video games through language-guided policy generation. By tran…
Pre-Trained Video Generative Models as World Simulators
Haoran He, Yang Zhang, Liang Lin +2
Video generative models pre-trained on large-scale internet datasets have achieved remarkable success, excelling at producing realistic synthetic videos. However, they often genera…