11 papers
Personalizing MLLMs via Reinforced Multimodal Reference Game
Deepayan Das, Davide Talon, Yiming Wang +2
Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Although prior work has shown t…
Is SAM3 ready for pathology segmentation?
Qiuyu Kong, Shakiba Sharifi, Yiming Wang +2
Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales, where traditional methods…
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
Huilai Li, Xiaomeng Di, Ying Xing +3
Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Fac…
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
Yiming Wang, Frederick W. B. Li, Jingyun Wang
Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP with disentangled embeddings an…
Specificity-aware reinforcement learning for fine-grained open-world classification
Samuele Angheben, Davide Berasi, Alessandro Conti +2
Classifying fine-grained visual concepts under open-world settings, i.e., without a predefined label set, demands models to be both accurate and specific. Recent reasoning Large Mu…
Training-free Online Video Step Grounding
Luca Zanella, Massimiliano Mancini, Yiming Wang +2
Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled tr…