37 papers
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challengi…
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling
Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7
CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +36
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
Jaehyun Jang, Eunseop Yoon, Hee Suk Yoon +3
While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying solely on output token distributio…
Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?
Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson +2
Synthetic accented speech is a promising way to improve automatic speech recognition (ASR) when real accented recordings are scarce. We ask what makes such data useful for ASR fine…
Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon +2
Recent flow-matching text-to-speech (TTS) models, such as F5-TTS, rely on a reference transcript at inference time, obtained from an external ASR system. This dependency makes zero…