activity
20222025
most citedCaption Anything: Interactive Image Description with Diverse Multimodal Controls

19 citations · 43 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

22 papers · 1 filter

cs.CV2025

When to Think and When to Look: Uncertainty-Guided Lookback

Jing Bi, Filippos Bellos, Junjia Guo +8

Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large…

cs.CV2025

FreSca: Scaling in Frequency Space Enhances Diffusion Models

Chao Huang, Susan Liang, Yunlong Tang +4

Latent diffusion models (LDMs) have achieved remarkable success in a variety of image tasks, yet achieving fine-grained, disentangled control over global structures versus fine det…

cs.CV2025

MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

Yolo Y. Tang, Pinxin Liu, Zhangyun Tan +11

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains uncle…

cs.CV2025

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

Yunlong Tang, Jing Bi, Chao Huang +16

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects…

cs.CV2025

Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability

Jiani Liu, Zhiyuan Wang, Zeliang Zhang +4

Vision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. However, their unique architectural proper…

cs.CV2025

VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity

Jing Bi, Junjia Guo, Susan Liang +8

Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLL…