activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Qing'an Liu, Juntong Feng, Yuhao Wang +6

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pur…

cs.CV2025

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

Harold Haodong Chen, Disen Lan, Wen-Jie Shu +10

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consi…

cs.CV2025

Interleaving Reasoning for Better Text-to-Image Generation

Wenxuan Huang, Shuang Chen, Zheyong Xie +15

Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction followin…

cs.CV2025

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Jihai Zhang, Tianle Li, Linjie Li +2

Recent advancements in unified vision-language models (VLMs), which integrate both visual understanding and generation capabilities, have attracted significant attention. The under…

cs.CV2025

Visually Interpretable Subtask Reasoning for Visual Question Answering

Yu Cheng, Arushi Goel, Hakan Bilen

Answering complex visual questions like `Which red furniture can be used for sitting?' requires multi-step reasoning, including object recognition, attribute filtering, and relatio…

cs.CV2025

BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries

Tianle Li, Yongming Rao, Winston Hu +1

Encoder-free multimodal large language models(MLLMs) eliminate the need for a well-trained vision encoder by directly processing image tokens before the language model. While this…