3 citations · 3 across the 3 of their papers we have counts for
5 papers
A Robust framework for sound event localization and detection on real recordings
Jin Sob Kim, Hyun Joon Park, Wooseok Shin +1
This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound…
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
Jin Sob Kim, Hyun Joon Park, Wooseok Shin +1
Recent speaker verification studies have achieved notable success by leveraging layer-wise output from pre-trained Transformer models. However, few have explored the advancements i…
Leveraging Out-of-Distribution Unlabeled Images: Semi-Supervised Semantic Segmentation with an Open-Vocabulary Model
Wooseok Shin, Jisu Kang, Hyeonki Jeong +2
In semi-supervised semantic segmentation, existing studies have shown promising results in academic settings with controlled splits of benchmark datasets. However, the potential be…
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
Hyun Joon Park, Jeongmin Liu, Jin Sob Kim +3
We introduce RapFlow-TTS, a rapid and high-fidelity TTS acoustic model that leverages velocity consistency constraints in flow matching (FM) training. Although ordinary differentia…
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
Byung Hoon Lee, Wooseok Shin, Sung Won Han
The word-level lipreading approach typically employs a two-stage framework with separate frontend and backend architectures to model dynamic lip movements. Each component has been…