12 papers
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Bohai Gu, Yueyang Yuan, Taiyi Wu +9
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learnin…
Decomposing and Measuring Evaluation Awareness
Changling Li, Terry Jingchen Zhang, Jie Zhang +3
Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a…
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
Bohai Gu, Taiyi Wu, Yueyang Yuan +9
Recent video-based world models have made pixel-space environments interactive at the camera level: users can navigate viewpoints while the model generates coherent visual continua…
Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
Lukas Twist, Helen Yannakoudakis, Jie M. Zhang
Explicit reasoning models are trained to produce intermediate reasoning traces before final answers, but downstream fine-tuning is often performed on ordinary instruction-response…
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
Qiran Zhang, Yuheng Wang, Runde Yang +9
Programmatic video generation through code offers geometric precision and temporal coherence beyond pixel-level diffusion models, yet rigorously evaluating whether language models…
Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation
Shuyin Ouyang, Zhaozhi Qian, Faroq AL-Tam +2
Reinforcement Learning (RL) is an important paradigm for aligning Diffusion Language Models (DLMs) toward functional correctness in code generation. However, these models often enc…