4 papers
ENCORE: Entropy-Guided Cropping and Attention Regularization for Robust Vision--Language Understanding
Yuanhao Sun, Huawei Ji, Jiaxin Ding +2
Vision-Language Models (VLMs) perform well on diverse vision-language tasks, but transformer-based visual encoders split images into fixed-resolution sub-images, compromising objec…
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models
Huawei Ji, Yuanhao Sun, Yuan Jin +4
Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vision-language models (VLMs).…
One-shot Compositional 3D Head Avatars with Deformable Hair
Yuan Sun, Xuan Wang, WeiLi Zhang +3
We propose a compositional method for constructing a complete 3D head avatar from a single image. Prior one-shot holistic approaches frequently fail to produce realistic hair dynam…
DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models
Yangyang Qian, Yuan Sun, Yu Guo
Generating and editing dynamic 3D head avatars are crucial tasks in virtual reality and film production. However, existing methods often suffer from facial distortions, inaccurate…