activity
20182026
most citedEAGLE: Egocentric AGgregated Language-video Engine

3 citations · 5 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

23 papers · 1 filter

cs.CV2026

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models main…

cs.CV2026

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

Ce Zhang, Jing Bi, Jinxi He +9

Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a…

cs.CV2026

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents

Susan Liang, Chao Huang, Filippos Bellos +3

Active visual agents solve fine-grained image tasks by interleaving reasoning with image-grounding actions across multiple turns. However, deployment-time rollout budgets are rarel…

cs.CV20261 cited

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang +8

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types,…

cs.CV2026

TDMM-LM: Bridging Facial Understanding and Animation via Language Models

Luchuan Song, Pinxin Liu, Haiyang Liu +7

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage f…

cs.CV2026

Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?

Susan Liang, Chao Huang, Filippos Bellos +7

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new…