9 papers
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
Yidong Huang, Zun Wang, Han Lin +6
Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general…
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation
Jacob Levy, Tyler Westenbroek, Kevin Huang +6
Robot learning requires adaptation methods that improve reliably from limited, mixed-quality interaction data. This is especially challenging in long-horizon, contact-rich tasks, w…
Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
Aileen Liao, Dong-Ki Kim, Max Olan Smith +2
As a robot senses and selects actions, the world keeps changing. This inference delay creates a gap of tens to hundreds of milliseconds between the observed state and the state at…
GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
Haoyang He, Jay Patrikar, Dong-Ki Kim +5
Recent advances in video world modeling have enabled large-scale generative models to simulate embodied environments with high visual fidelity, providing strong priors for predicti…
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
Yidong Huang, Zun Wang, Han Lin +5
Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However…
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
Jason Jabbour, Dong-Ki Kim, Max Smith +6
Vision-Language-Action (VLA) models have advanced robotic capabilities but remain challenging to deploy on resource-limited hardware. Pruning has enabled efficient compression of l…