works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

Wenzhuo Zhao, Xiuzhi Li, Zhongkuan Mao +6

The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conv…

cs.CV2026

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9

The paper proposes a training‑free, single‑pass method that routes intermediate‑layer visual evidence to improve high‑resolution visual question answering without extra image proce…

cs.CV2026

Camouflage-aware Image-Text Retrieval via Expert Collaboration

Yao Jiang, Zhongkuan Mao, Xuan Wu +2

Camouflaged scene understanding (CSU) has attracted significant attention due to its broad practical implications. However, in this field, robust image-text cross-modal alignment r…

cs.CV2024

Effectiveness Assessment of Recent Large Vision-Language Models

Yao Jiang, Xinyu Yan, Ge-Peng Ji +5

The advent of large vision-language models (LVLMs) represents a remarkable advance in the quest for artificial general intelligence. However, the model's effectiveness in both spec…

cs.CV2024

Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models

Dehong Kong, Siyuan Liang, Xiaopeng Zhu +2

Visual language pre-training (VLP) models have demonstrated significant success across various domains, yet they remain vulnerable to adversarial attacks. Addressing these adversar…