5 papers
SelFusion: Self-distillation for Diffusion Language Models
Hyeongsoo Lim, Jinyoung Kim, Eunseo Seo +2
Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) large language models (LLMs), but their degraded generation quality limits practic…
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo +4
Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English langua…
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
Ju Yeon Kang, Jaehong Park, Semin Kim +2
Recently, AI-generated image detection has gained increasing attention, as the rapid advancement of image generation technologies has raised serious concerns about their potential…
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
Ju Yeon Kang, Ji Won Yoon, Semin Kim +2
Recently, fake audio detection has gained significant attention, as advancements in speech synthesis and voice conversion have increased the vulnerability of automatic speaker veri…
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
Hyeonseung Lee, Ji Won Yoon, Sungsoo Kim +1
Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and…