5 papers · 1 filter
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
Hanpeng Liu, Yaqian Li, Zidan Wang +6
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organiz…
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
Hanpeng Liu, Yaqian Li, Zidan Wang +5
Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision…
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
Hanpeng Liu, Zidan Wang, Shuoxi Zhang +2
The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient processing of long sequence tasks…
Neural Collapse Inspired Knowledge Distillation
Shuoxi Zhang, Zijian Song, Kun He
Existing knowledge distillation (KD) methods have demonstrated their ability in achieving student network performance on par with their teachers. However, the knowledge gap between…
You Only Need Less Attention at Each Stage in Vision Transformers
Shuoxi Zhang, Hanpeng Liu, Stephen Lin +1
The advent of Vision Transformers (ViTs) marks a substantial paradigm shift in the realm of computer vision. ViTs capture the global information of images through self-attention mo…