4 papers
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
KiHyun Nam, Jungwoo Heo, Siu Bae +2
As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-sp…
SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim +4
Self-supervised learning (SSL) has pushed speaker verification accuracy close to state-of-the-art levels, but the Transformer backbones used in most SSL encoders hinder on-device a…
Token-based Attractors and Cross-attention in Spoof Diarization
Kyo-Won Koo, Chan-yeong Lim, Jee-weon Jung +2
Spoof diarization identifies ``what spoofed when" in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward thi…
SEED: Speaker Embedding Enhancement Diffusion Model
KiHyun Nam, Jungwoo Heo, Jee-weon Jung +4
A primary challenge when deploying speaker recognition systems in real-world applications is performance degradation caused by environmental mismatch. We propose a diffusion-based…