collaborators

5 papers

cs.CV2026

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

Yuqian Yuan, Wenqiao Zhang, Juekai Lin +7

Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level…

cs.CV2025

GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval

Bowen Yang, Yun Cao, Chen He +1

Text-to-video retrieval requires precise alignment between language and temporally rich audio-video signals. However, existing methods often emphasize visual cues while underutiliz…

cs.CV2025

MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention

Yuqi Pang, Bowen Yang, Yun Cao +3

Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual…

cs.CV2025

Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics

Yiran He, Yun Cao, Bowen Yang +1

The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs)…

cs.CV2025

Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning

Yuqi Pang, Bowen Yang, Haoqin Tu +2

Although Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large…