activity
20162024
most citedThe VoxCeleb Speaker Recognition Challenge: A Retrospective

22 citations · 61 across the 25 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2024

VoxSim: A perceptual voice similarity dataset

Junseok Ahn, Youkyum Kim, Yeunju Choi +4

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on…

eess.AS2024

Lightweight Audio Segmentation for Long-form Speech Translation

Jaesong Lee, Soyoon Kim, Hanbyul Kim +1

Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio m…

eess.AS2024

FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching

Chaeyoung Jung, Suyeon Lee, Ji-Hoon Kim +1

This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches hav…

eess.AS2024

FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder

Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang +2

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key componen…

eess.AS20231 cited

Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model

Suyeon Lee, Chaeyoung Jung, Youngjoon Jang +2

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their…

eess.AS2023

Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

Hee-Soo Heo, KiHyun Nam, Bong-Jin Lee +4

In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from spea…