collaborators

5 papers

cs.CV2026

Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation

Chongjie Ye, Cheng Cao, Chuanyu Pan +4

Recent multimodal large language models have achieved strong performance in unified text and image understanding and generation, yet extending such native capability to 3D remains…

cs.CV2025

LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models

Yiming Hao, Mutian Xu, Chongjie Ye +4

Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to th…

cs.AI2025

VC-Agent: An Interactive Agent for Customized Video Dataset Collection

Yidan Zhang, Mutian Xu, Yiming Hao +6

Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and tim…

cs.GR2025

IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects

Xiaokang Wei, Zizheng Yan, Zhangyang Xiong +3

Estimating albedo (a.k.a., intrinsic image decomposition) from single RGB images captured in real-world environments (e.g., the MVImgNet dataset) presents a significant challenge d…

cs.CV2025

TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation

Hongxiang Zhao, Xingchen Liu, Mutian Xu +3

We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for ro…