collaborators

5 papers

cs.CV2026

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

Shibo Gao, Chongxiao Wang, Chenglong Huang +10

Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identit…

cs.CV2026

MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs

Yilian Liu, Xiaojun Jia, Guoshun Nan +6

Multimodal Large Language Models (MLLMs) have achieved remarkable performance but remain vulnerable to jailbreak attacks that can induce harmful content and undermine their secure…

cs.CV2026

ConFoThinking: Consolidated Focused Attention Driven Thinking for Visual Question Answering

Zhaodong Wu, Haochen Xue, Qi Cao +5

Thinking with Images improves fine-grained VQA for MLLMs by emphasizing visual cues. However, tool-augmented methods depend on the capacity of grounding, which remains unreliable f…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…

cs.CV2025

AdsQA: Towards Advertisement Video Understanding

Xinwei Long, Kai Tian, Peng Xu +10

Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…