3 papers
cs.CV2026
ASD: Multi-Level Consistency-Driven Representation Learning
Jin Hong, Jisoo Park, Junseok Kwon
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrad…
cs.CV2025
Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
Jisoo Park, Seonghak Lee, Guisik Kim +2
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background…
cs.CV2025
Real-Aware Residual Model Merging for Deepfake Detection
Jinhee Park, Guisik Kim, Choongsang Cho +1
Deepfake generators evolve quickly, making exhaustive data collection and repeated retraining impractical. We argue that model merging is a natural fit for deepfake detection: unli…