Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
ASD: Multi-Level Consistency-Driven Representation Learning
Jin Hong, Jisoo Park, Junseok Kwon
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrad…
cs.CV2025
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
Jisoo Park, Seonghak Lee, Guisik Kim +2
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background…
cs.CV2025
RM: Real-Aware Residual Model Merging for Robust and Generalizable Deepfake Detection
Jinhee Park, Guisik Kim, Choongsang Cho +1
Deepfake generators evolve rapidly, making exhaustive data collection and repeated retraining impractical. Unlike generic multi-task settings, deepfake specialists share a common b…