collaborators

14 papers

cs.CV2026

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents

Susan Liang, Chao Huang, Filippos Bellos +3

Active visual agents solve fine-grained image tasks by interleaving reasoning with image-grounding actions across multiple turns. However, deployment-time rollout budgets are rarel…

cs.CV2026

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang +8

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types,…

cs.CV2025

Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination

Yolo Y. Tang, Daiki Shimada, Hang Hua +4

Understanding text-rich videos requires reading small, transient textual cues that often demand repeated inspection. Yet most video QA models rely on single-pass perception over fi…

cs.CV2025

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

Yolo Y. Tang, Jing Bi, Pinxin Liu +24

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…

cs.CL2025

Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)

Jing Bi, Susan Liang, Xiaofei Zhou +16

Reasoning is central to human intelligence, enabling structured problem-solving across diverse tasks. Recent advances in large language models (LLMs) have greatly enhanced their re…

cs.CV2025

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang +3

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: ho…