8 papers
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
Minghao Tian, Yunfei Xie, Chen Wei
Group Relative Policy Optimization (GRPO) has been a key driver of recent progress in reinforcement learning with verifiable rewards (RLVR) for large language models, but it is typ…
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
Sinin Zhang, Yunfei Xie, Yuxuan Cheng +2
Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios th…
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Kyuyoung Kim, Kevin Wang, Yunfei Xie +7
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifiable rewards typically optimizes only fina…
MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games
Yunfei Xie, Kevin Wang, Bobby Cheng +9
Multi-turn, multi-agent LLM game evaluations often exhibit substantial run-to-run variance. In long-horizon interactions, small early deviations compound across turns and are ampli…
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie +7
This paper introduces Story-Iter, a new training-free iterative paradigm to enhance long-story generation. Unlike existing methods that rely on fixed reference images to construct…
Play to Generalize: Learning to Reason Through Game Play
Yunfei Xie, Yinsong Ma, Shiyi Lan +3
Developing reasoning capabilities in multimodal large language models (MLLMs) remains challenging. Motivated by literature suggesting that gameplay promotes transferable reasoning…