4 papers
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
Bowen Yang, Yun Cao, Chen He +1
Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutiliz…
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
Yuqi Pang, Bowen Yang, Yun Cao +3
Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual…
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
Yiran He, Yun Cao, Bowen Yang +1
The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs)…
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
Yuqi Pang, Bowen Yang, Haoqin Tu +2
Although Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large…