collaborators

11 papers

cs.CV2026

Personalizing MLLMs via Reinforced Multimodal Reference Game

Deepayan Das, Davide Talon, Yiming Wang +2

Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Although prior work has shown t…

cs.CV2026

Is SAM3 ready for pathology segmentation?

Qiuyu Kong, Shakiba Sharifi, Yiming Wang +2

Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales, where traditional methods…

cs.CV2026

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

Huilai Li, Xiaomeng Di, Ying Xing +3

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Fac…

cs.CV2026

Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition

Yiming Wang, Frederick W. B. Li, Jingyun Wang

Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP with disentangled embeddings an…

cs.CV2026

Specificity-aware reinforcement learning for fine-grained open-world classification

Samuele Angheben, Davide Berasi, Alessandro Conti +2

Classifying fine-grained visual concepts under open-world settings, i.e., without a predefined label set, demands models to be both accurate and specific. Recent reasoning Large Mu…

cs.CV2025

Training-free Online Video Step Grounding

Luca Zanella, Massimiliano Mancini, Yiming Wang +2

Given a task and a set of steps composing it, Video Step Grounding (VSG) aims to detect which steps are performed in a video. Standard approaches for this task require a labeled tr…