5 papers
Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs
Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera +2
Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability…
Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
Chao Wang, Chengan Che, Xinyue Chen +2
Counterfactual explanations (CFEs) are minimal and semantically meaningful modifications of the input of a model that alter the model predictions. They highlight the decisive featu…
A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Ranking
Chengan Che, Chao Wang, Xinyue Chen +2
Procedural activities, ranging from routine cooking to complex surgical operations, are highly structured sequences of actions performed in a specific temporal order. Despite the s…
LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings
Chengan Che, Chao Wang, Tom Vercauteren +2
Traditional open-access datasets focusing on surgical procedures are often limited by their small size, typically consisting of fewer than 100 videos and less than 30 hours of foot…
OwMatch: Conditional Self-Labeling with Consistency for Open-World Semi-Supervised Learning
Shengjie Niu, Lifan Lin, Jian Huang +1
Semi-supervised learning (SSL) offers a robust framework for harnessing the potential of unannotated data. Traditionally, SSL mandates that all classes possess labeled instances. H…