9 papers
Advances in 4D Representation: Geometry, Motion, and Interaction
Mingrui Zhao, Sauradip Nag, Kai Wang +5
We present a survey on 4D generation and reconstruction, a fast-evolving subfield of computer graphics whose developments have been propelled by recent advances in neural fields, g…
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
Aditya Vora, Sauradip Nag, Kai Wang +1
We present ATOP (Articulate That Object Part), a novel few-shot method based on motion personalization to articulate a static 3D object with respect to a part and its motion as pre…
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
Jiachun Jin, Zetong Zhou, Xiao Yang +4
Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
Hao Zhang, Lue Fan, Weikang Bian +4
We present ReinDriveGen, a framework that enables full controllability over dynamic driving scenes, allowing users to freely edit actor trajectories to simulate safety-critical cor…
Hierarchical Transformers for Unsupervised 3D Shape Abstraction
Aditya Vora, Lily Goli, Andrea Tagliasacchi +1
We introduce HiT, a novel hierarchical neural field representation for 3D shapes that learns general hierarchies in a coarse-to-fine manner across different shape categories in an…
VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
Hui Zhou, Siyuan Huang, Minxing Li +3
Vision Language Action models have significantly advanced general purpose robotic manipulation by harnessing large scale pretrained vision and language representations. Among exist…