4 papers
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu +8
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Bohai Gu, Yueyang Yuan, Taiyi Wu +9
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learnin…
Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation
Zhen Zhou, Jian Liu, Biwen Lei +10
Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typic…
Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation
Zichen Geng, Zeeshan Hayder, Bo Miao +3
Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compre…