9 papers
NeoMap: Training-free Novel-View Synthesis from Single Images and Videos
Jinxi Li, Tianyi Zhang, Yafei Yang +4
We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assumption that pre-trained video m…
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
Peng Yun, Shouwang Huang, Hao Li +3
Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action models and world models strugg…
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation
Zihui Zhang, Zhixuan Sun, Yafei Yang +3
We address the challenging task of 3D object segmentation in complex scene point clouds without relying on any scene-level human annotations during training. Existing methods are t…
EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
Jiahao Chen, Zihui Zhang, Yafei Yang +4
We introduce EvObj for unsupervised 3D instance segmentation that bridges the geometric domain gap between synthetic pretraining data and real-world point clouds. Current methods s…
PhysInOne: Visual Physics Learning and Reasoning in One Suite
Siyuan Zhou, Hejun Wang, Hu Cheng +36
We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to mere…
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
Rong Zhang, Jinxiao Li, Jingnan Wang +6
Garment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practi…