3 papers
eess.AS2026
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
Helin Wang, Bowen Shi, Andros Tjandra +6
The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on g…
eess.AS2025
SAM Audio: Segment Anything in Audio
Bowen Shi, Andros Tjandra, John Hoffman +11
General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separ…
cs.CV2025
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
Jiahe Zhao, Rongkun Zheng, Yi Wang +2
In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input…