collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt +2

Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and ver…

cs.CV2026

Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models

Dhruba Ghosh, Yuhui Zhang, Ludwig Schmidt

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and mul…

cs.CV2025

Data or Language Supervision: What Makes CLIP Better than DINO?

Yiming Liu, Yuhui Zhang, Dhruba Ghosh +2

CLIP outperforms self-supervised models like DINO as vision encoders for vision-language models (VLMs), but it remains unclear whether this advantage stems from CLIP's language sup…

cs.CV2025

Closing the Modality Gap for Mixed Modality Search

Binxu Li, Yuhui Zhang, Xiaohan Wang +3

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world ap…

cs.CV2025

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Yuhui Zhang, Yuchang Su, Yiming Liu +9

The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-en…