12 papers
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Hongyu Li, Wanjia Fu, Xiaoyan Cong +11
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to…
Turbo-GS: Accelerating 3D Gaussian Fitting for High-Quality Radiance Fields
Ankit Dhiman, Tao Lu, R Srinath +5
Novel-view synthesis plays a crucial role in computer vision with applications in 3D reconstruction, mixed reality, and robotics. Recent approaches, such as 3D Gaussian Splatting (…
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
Zekun Li, Sizhe An, Chengcheng Tang +7
Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language ge…
GenHSI: Controllable Generation of Human-Scene Interaction Videos
Zekun Li, Rui Zhou, Rahul Sajnani +3
Large-scale pre-trained video diffusion models have exhibited remarkable capabilities in diverse video generation. However, existing solutions face several challenges in generating…
DyTact: Capturing Dynamic Contacts in Hand-Object Manipulation
Xiaoyan Cong, Angela Xing, Chandradeep Pokhariya +2
Reconstructing dynamic hand-object contacts is essential for realistic manipulation in AI character animation, XR, and robotics, yet it remains challenging due to heavy occlusions,…
Art3D: Training-Free 3D Generation from Flat-Colored Illustration
Xiaoyan Cong, Jiayi Shen, Zekun Li +3
Large-scale pre-trained image-to-3D generative models have exhibited remarkable capabilities in diverse shape generations. However, most of them struggle to synthesize plausible 3D…