12 papers
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Yiyang Cai, Nan Chen, Rongchang Xie +8
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most app…
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
Nan Chen, Yiyang Cai, Rongchang Xie +7
Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two scenarios: in-domain, which r…
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Zhangzhao Liang, Xiaofen Xing, Mingyue Yang +2
Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment constraints. Existing co-speech generatio…
Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning
Huayi Zhou, Wei Gao, Dekun Lu +12
End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However,…
Masked Representation Modeling for Domain-Adaptive Segmentation
Wenlve Zhou, Zhiheng Zhou, Tiantao Xian +3
Unsupervised domain adaptation (UDA) for semantic segmentation seeks to transfer models from a labeled source domain to an unlabeled target domain. While auxiliary self-supervised…
Pano360: Perspective to Panoramic Vision with Geometric Consistency
Zhengdong Zhu, Weiyi Xue, Zuyuan Yang +2
Prior panorama stitching approaches heavily rely on pairwise feature correspondences and are unable to leverage geometric consistency across multiple views. This leads to severe di…