3 papers
cs.CV2026
LIVE: Long-horizon Interactive Video World Modeling
Junchao Huang, Ziyang Ye, Xinting Hu +5
Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon gene…
cs.LG2026
Self-Hinting Language Models Enhance Reinforcement Learning
Baohao Liao, Hanze Dong, Xinxing Xu +2
Group Relative Policy Optimization (GRPO) has recently emerged as a practical recipe for aligning large language models with verifiable objectives. However, under sparse terminal r…
cs.AI2026
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
Haitian Zhong, Jixiu Zhai, Lei Song +3
Multi-turn tool calling is challenging for Large Language Models (LLMs) because rewards are sparse and exploration is expensive. A common recipe, SFT followed by GRPO, can stall wh…