collaborators

17 papers

cs.CV2026

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents

Susan Liang, Chao Huang, Filippos Bellos +3

Active visual agents solve fine-grained image tasks by interleaving reasoning with image-grounding actions across multiple turns. However, deployment-time rollout budgets are rarel…

cs.CV2026

When to Think and When to Look: Uncertainty-Guided Lookback

Jing Bi, Filippos Bellos, Junjia Guo +8

Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large…

cs.CV2026

Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?

Susan Liang, Chao Huang, Filippos Bellos +7

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new…

cs.SD2026

Semantic visually-guided acoustic highlighting with large vision-language models

Junhua Huang, Chao Huang, Chenliang Xu

Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-inten…

cs.CV2025

Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination

Yolo Y. Tang, Daiki Shimada, Hang Hua +4

Understanding text-rich videos requires reading small, transient textual cues that often demand repeated inspection. Yet most video QA models rely on single-pass perception over fi…

cs.CV2025

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

Yolo Y. Tang, Jing Bi, Pinxin Liu +24

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…