activity
20242026
collaborators

10 papers

cs.CV2026

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy +5

Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic s…

cs.CV2025

Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning

Shengguang Wu, Xiaohan Wang, Yuhui Zhang +2

Spatial reasoning in 3D scenes requires precise geometric calculations that challenge vision-language models. Visual programming addresses this by decomposing problems into steps c…

cs.CV2025

Closing the Modality Gap for Mixed Modality Search

Binxu Li, Yuhui Zhang, Xiaohan Wang +3

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world ap…

cs.CL2025

A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI

Alejandro Lozano, Min Woo Sun, James Burgess +16

Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bot…

cs.CV2025

Video Action Differencing

James Burgess, Xiaohan Wang, Yuhui Zhang +5

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences betw…

cs.CV2025

SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection

Devanish N. Kamtam, Joseph B. Shrager, Satya Deepya Malla +5

Background: We evaluate SAM 2 for surgical scene understanding by examining its semantic segmentation capabilities for organs/tissues both in zero-shot scenarios and after fine-tun…