14 papers
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
Yidong Huang, Zun Wang, Han Lin +6
Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general…
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
Yutian Chen, Yuheng Qiu, Ruogu Li +4
We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a l…
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation
Jacob Levy, Tyler Westenbroek, Kevin Huang +6
Robot learning requires adaptation methods that improve reliably from limited, mixed-quality interaction data. This is especially challenging in long-horizon, contact-rich tasks, w…
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
Aileen Liao, Dong-Ki Kim, Max Olan Smith +2
As a robot senses and selects actions, the world keeps changing. This inference delay creates a gap of tens to hundreds of milliseconds between the observed state and the state at…
World Model Failure Classification and Anomaly Detection for Autonomous Inspection
Michelle Ho, Muhammad Fadhil Ginting, Isaac R. Ward +4
Autonomous inspection robots for monitoring industrial sites can reduce costs and risks associated with human-led inspection. However, accurate readings can be challenging due to o…
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
Haoyang He, Jay Patrikar, Dong-Ki Kim +5
Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for predicti…