collaborators

6 papers

cs.MM2026

Unveiling the Visual Counting Bottleneck in Vision-Language Models

Xingzhou Pang, Yifan Hou, Junling Wang +1

While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we…

cs.LG2026

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

Yijie Tong, Yifan Hou, Shaobo Cui +2

Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their application and benefits for vision-lan…

cs.CL2026

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

Yucheng Wang, Yifan Hou, Aydin Javadov +2

Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-modal reasoning remains underexplored,…

cs.CL2025

Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding

Ziheng Chi, Yifan Hou, Chenxi Pang +3

Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations…

cs.CL2025

Can Vision-Language Models Solve Visual Math Equations?

Monjoy Narayan Choudhury, Junling Wang, Yifan Hou +1

Despite strong performance in visual understanding and language-based reasoning, Vision-Language Models (VLMs) struggle with tasks requiring integrated perception and symbolic comp…

cs.CL2025

Do Vision-Language Models Really Understand Visual Language?

Yifan Hou, Buse Giledereli, Yilei Tu +1

Visual language is a system of communication that conveys information through symbols, shapes, and spatial arrangements. Diagrams are a typical example of a visual language depicti…