works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models

Botai Yuan, Yutian Zhou, Yingjie Wang +9

Recent benchmarks for medical Large Vision-Language Models (LVLMs) emphasize leaderboard accuracy, overlooking reliability and safety. We study sycophancy -- models' tendency to un…

cs.CV2025

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Rong-Cheng Tu, Wenhao Sun, Hanzhe You +4

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a compositional query, consisting of a reference image and a modifying text-without relying on anno…

cs.CV2025

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval

Rong-Cheng Tu, Zhao Jin, Jingyi Liao +4

Existing Zero-Shot Composed Image Retrieval (ZS-CIR) methods typically train adapters that convert reference images into pseudo-text tokens, which are concatenated with the modifyi…

cs.CV2025

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Huanjin Yao, Qixiang Yin, Jingyi Zhang +8

In this work, we aim to incentivize the reasoning ability of Multimodal Large Language Models (MLLMs) via reinforcement learning (RL) and develop an effective approach that mitigat…

cs.CV2024

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Huanjin Yao, Jiaxing Huang, Wenhao Wu +8

In this work, we aim to develop an MLLM that understands and solves questions by learning to create each intermediate step of the reasoning involved till the final answer. To this…