3 papers
cs.CV2026
Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
Ziang Wu, Peng Jin, Qishen Yin +4
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the stan…
cs.CV2024
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
Peng Jin, Hao Li, Li Yuan +2
Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representatio…
cs.CV2024
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Guowei Xu, Peng Jin, Ziang Wu +4
Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and…