4 papers
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
Hanpeng Liu, Yaqian Li, Zidan Wang +6
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organiz…
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
Hanpeng Liu, Yaqian Li, Zidan Wang +5
Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision…
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
Hanpeng Liu, Zidan Wang, Shuoxi Zhang +2
The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting significant attention due to its efficient processing of long sequence tasks…
Enhancing Pre-Trained Model-Based Class-Incremental Learning through Neural Collapse
Kun He, Zijian Song, Shuoxi Zhang +1
Class-Incremental Learning (CIL) is a critical capability for real-world applications, enabling learning systems to adapt to new tasks while retaining knowledge from previous ones.…