5 citations · 28 across the 16 of their papers we have counts for
20 papers · 1 filter
Reinforced Disentanglement for Face Swapping without Skip Connection
Xiaohang Ren, Xingyu Chen, Pengfei Yao +2
The SOTA face swap models still suffer the problem of either target identity (i.e., shape) being leaked or the target non-identity attributes (i.e., background, hair) failing to be…
Visual Instruction Tuning with Polite Flamingo
Delong Chen, Jianfeng Liu, Wenliang Dai +1
Recent research has demonstrated that the multi-task fine-tuning of multi-modal Large Language Models (LLMs) using an assortment of annotated downstream vision-language datasets si…
ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers
Jingfeng Yao, Xinggang Wang, Shusheng Yang +1
Recently, plain vision Transformers (ViTs) have shown impressive performance on various computer vision tasks, thanks to their strong modeling capacity and large-scale pretraining.…
FashionTex: Controllable Virtual Try-on with Text and Texture
Anran Lin, Nanxuan Zhao, Shuliang Ning +3
Virtual try-on attracts increasing research attention as a promising way for enhancing the user experience for online cloth shopping. Though existing methods can generate impressiv…
An Effective Motion-Centric Paradigm for 3D Single Object Tracking in Point Clouds
Chaoda Zheng, Xu Yan, Haiming Zhang +4
3D single object tracking in LiDAR point clouds (LiDAR SOT) plays a crucial role in autonomous driving. Current approaches all follow the Siamese paradigm based on appearance match…
Mimic3D: Thriving 3D-Aware GANs via 3D-to-2D Imitation
Xingyu Chen, Yu Deng, Baoyuan Wang
Generating images with both photorealism and multiview 3D consistency is crucial for 3D-aware GANs, yet existing methods struggle to achieve them simultaneously. Improving the phot…