11 papers
DirectTryOn: One-Step Virtual Try-On via Straightened Conditional Transport
Xianbing Sun, Jiahui Zhan, Liqing Zhang +1
Recent diffusion- and flow-based VTON methods achieve strong results with pretrained generative models, but their reliance on multi-step sampling incurs high inference cost, while…
Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation
Xinhao Hu, Yiyi Zhang, Liqing Zhang +1
Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human…
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
Yikun Ji, Yan Hong, Bowen Deng +5
The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs…
Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
Yujie Gao, Yao Xiao, Xiangnan Zhu +4
Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-fidelity pipelines often rely…
AVGGT: Rethinking Global Attention for Accelerating VGGT
Xianbing Sun, Zhikai Zhu, Zhengyu Lou +5
Models such as VGGT and have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-at…
Towards Source-Aware Object Swapping with Initial Noise Perturbation
Jiahui Zhan, Xianbing Sun, Xiangnan Zhu +4
Object swapping aims to replace a source object in a scene with a reference object while preserving object fidelity, scene fidelity, and object-scene harmony. Existing methods eith…