activity
20202025
most citedAnti-Spoofing Using Transfer Learning with Variational Information Bottleneck

22 citations · 35 across the 14 of their papers we have counts for

collaborators
Showing eess.ASShow all

13 papers · 1 filter

eess.AS2025

Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

Sungnyun Kim, Kangwook Jang, Sungwoo Cho +3

This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directl…

eess.AS2025

ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction

Minu Kim, Kangwook Jang, Hoirin Kim

Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on im…

eess.AS2025

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

Minu Kim, Kangwook Jang, Hoirin Kim

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language s…

eess.AS2024

Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition

Sungnyun Kim, Kangwook Jang, Sangmin Bae +2

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of vide…

eess.AS2024

One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection

Hyun Myung Kim, Kangwook Jang, Hoirin Kim

As speech synthesis systems continue to make remarkable advances in recent years, the importance of robust deepfake detection systems that perform well in unseen systems has grown.…

eess.AS2023★ 6 cited

Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation

Kangwook Jang, Sungnyun Kim, Se-Young Yun +1

Transformer-based speech self-supervised learning (SSL) models, such as HuBERT, show surprising performance in various speech processing tasks. However, huge number of parameters i…