2 citations · 4 across the 4 of their papers we have counts for
6 papers
Look Who's Talking: Active Speaker Detection in the Wild
You Jin Kim, Hee-Soo Heo, Soyeon Choe +5
In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audi…
Overcoming label noise in audio event detection using sequential labeling
Jae-Bin Kim, Seongkyu Mun, Myungwoo Oh +3
This paper addresses the noisy label issue in audio event detection (AED) by refining strong labels as sequential labels with inaccurate timestamps removed. In AED, strong labels c…
FaceFilter: Audio-visual speech separation using still images
Soo-Whan Chung, Soyeon Choe, Joon Son Chung +1
The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that…
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun +7
The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level repre…
The sound of my voice: speaker representation loss for target voice separation
Seongkyu Mun, Soyeon Choe, Jaesung Huh +1
Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for au…
Orthonormal Embedding-based Deep Clustering for Single-channel Speech Separation
Soyeon Choe, Soo-Whan Chung, Youna Ji +1
Deep clustering is a deep neural network-based speech separation algorithm that first trains the mixed component of signals with high-dimensional embeddings, and then uses a cluste…