4 papers
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
Songen Gu, Yunuo Cai, Tianyu Wang +2
Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors…
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
Binjie Liu, Lina Liu, Sanyi Zhang +5
This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation…
VRsketch2Gaussian: 3D VR Sketch Guided 3D Object Generation with Gaussian Splatting
Songen Gu, Haoxuan Song, Binjie Liu +5
We propose VRSketch2Gaussian, a first VR sketch-guided, multi-modal, native 3D object generation framework that incorporates a 3D Gaussian Splatting representation. As part of our…