Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ASD: Multi-Level Consistency-Driven Representation Learning
Jin Hong, Jisoo Park, Junseok Kwon
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrad…
cs.CV2025
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
Jisoo Park, Seonghak Lee, Guisik Kim +2
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background…