6 papers
PointAction: 3D Points as Universal Action Representations for Robot Control
Mutian Tong, Han Jiang, Qiao Feng +2
Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation. How…
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
Yunzhou Song, Long Le, Yong-Hyun Park +7
Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on mo…
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
Chen Wang, Chuhao Chen, Yiming Huang +4
Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limit…
PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction
Qiao Feng, Yiming Huang, Yufu Wang +2
Reconstructing physically plausible human motion from monocular videos remains a challenging problem in computer vision and graphics. Existing methods primarily focus on kinematics…
Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation
Chuhao Chen, Zhiyang Dou, Chen Wang +5
Faithfully reconstructing textured shapes and physical properties from videos presents an intriguing yet challenging problem. Significant efforts have been dedicated to advancing s…
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
Xuyi Meng, Chen Wang, Jiahui Lei +3
Recent advances in 2D image generation have achieved remarkable quality,largely driven by the capacity of diffusion models and the availability of large-scale datasets. However, di…