2 papers
cs.CV2026
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
Hanpeng Liu, Yaqian Li, Zidan Wang +6
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organiz…
cs.CV2026
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
Hanpeng Liu, Yaqian Li, Zidan Wang +5
Despite the success of Large Vision--Language Models (LVLMs), most existing architectures suffer from a representation bottleneck: they rely on static, instruction-agnostic vision…