7 papers
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
Hanpeng Liu, Yaqian Li, Zidan Wang +6
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organiz…
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
Hanpeng Liu, Yaqian Li, Zidan Wang +5
Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision…
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
Hanpeng Liu, Zidan Wang, Shuoxi Zhang +2
The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient processing of long sequence tasks…
Enhancing Pre-Trained Model-Based Class-Incremental Learning through Neural Collapse
Kun He, Zijian Song, Shuoxi Zhang +1
Class-Incremental Learning (CIL) is a critical capability for real-world applications, enabling learning systems to adapt to new tasks while retaining knowledge from previous ones.…
Neural Collapse Inspired Knowledge Distillation
Shuoxi Zhang, Zijian Song, Kun He
Existing knowledge distillation (KD) methods have demonstrated their ability in achieving student network performance on par with their teachers. However, the knowledge gap between…
Siamese Transformer Networks for Few-shot Image Classification
Weihao Jiang, Shuoxi Zhang, Kun He
Humans exhibit remarkable proficiency in visual classification tasks, accurately recognizing and classifying new images with minimal examples. This ability is attributed to their c…