collaborators

8 papers

cs.CV2026

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera +2

Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability…

cs.CV2026

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations

Chao Wang, Chengan Che, Xinyue Chen +2

Counterfactual explanations (CFEs) are minimal and semantically meaningful modifications of the input of a model that alter the model predictions. They highlight the decisive featu…

cs.CV2026

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?

Chengan Che, Chao Wang, Jiayuan Huang +2

Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations t…

cs.CV2026

A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking

Chengan Che, Chao Wang, Xinyue Chen +2

Procedural activities, ranging from routine cooking to complex surgical operations, are highly structured sequences of actions performed in a specific temporal order. Despite the s…

cs.CV2026

LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings

Chengan Che, Chao Wang, Tom Vercauteren +2

Traditional open-access datasets focusing on surgical procedures are often limited by their small size, typically consisting of fewer than 100 videos and less than 30 hours of foot…

cs.CV2026

Average Calibration Losses for Reliable Uncertainty in Medical Image Segmentation

Theodore Barfoot, Luis C. Garcia-Peraza-Herrera, Samet Akcay +2

Deep neural networks for medical image segmentation are often overconfident, compromising both reliability and clinical utility. In this work, we propose differentiable formulation…