3 papers
cs.CV2026
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Lingxiao Li, Yifan Wang, Xinyan Gao +3
Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language mo…
cs.CV2026
Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning
Xinyan Gao, Haoran Hao, Xiangyu Yue
The rapid development of pretrained foundation models has enabled more general image segmentation. Multimodal large language models (MLLMs) have been widely explored for image segm…
cs.CL2025
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Yilei Jiang, Xinyan Gao, Tianshuo Peng +4
The integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language…