5 papers
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
Chongjie Ye, Cheng Cao, Chuanyu Pan +4
Recent multimodal large language models have achieved strong performance in unified text and image understanding and generation, yet extending such native capability to 3D remains…
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
Yiming Hao, Mutian Xu, Chongjie Ye +4
Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to th…
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
Yidan Zhang, Mutian Xu, Yiming Hao +6
Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and tim…
IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects
Xiaokang Wei, Zizheng Yan, Zhangyang Xiong +3
Estimating albedo (a.k.a., intrinsic image decomposition) from single RGB images captured in real-world environments (e.g., the MVImgNet dataset) presents a significant challenge d…
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
Hongxiang Zhao, Xingchen Liu, Mutian Xu +3
We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for ro…