activity
20232026
most citedVideoAgent: Long-form Video Understanding with Large Language Model as Agent

1 citations · 3 across the 15 of their papers we have counts for

collaborators
Showing 2025 · cs.CVShow all

7 papers · 2 filters

cs.CV2025

Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning

Shengguang Wu, Xiaohan Wang, Yuhui Zhang +2

Spatial reasoning in 3D scenes requires precise geometric calculations that challenge vision-language models. Visual programming addresses this by decomposing problems into steps c…

cs.CV2025

Closing the Modality Gap for Mixed Modality Search

Binxu Li, Yuhui Zhang, Xiaohan Wang +3

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world ap…

cs.CV2025

Video Action Differencing

James Burgess, Xiaohan Wang, Yuhui Zhang +5

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences betw…

cs.CV20251 cited

SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection

Devanish N. Kamtam, Joseph B. Shrager, Satya Deepya Malla +5

Background: We evaluate SAM 2 for surgical scene understanding by examining its semantic segmentation capabilities for organs/tissues both in zero-shot scenarios and after fine-tun…

cs.CV2025

Temporal Preference Optimization for Long-Form Video Understanding

Rui Li, Xiaohan Wang, Yuhui Zhang +3

Despite significant advancements in video large multimodal models (video-LMMs), achieving effective temporal grounding in long-form videos remains a challenge for existing models.…

cs.CV20251 cited

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

Alejandro Lozano, Min Woo Sun, James Burgess +13

The development of vision-language models (VLMs) is driven by large-scale and diverse multimodal datasets. However, progress toward generalist biomedical VLMs is limited by the lac…