activity
20242026
collaborators

37 papers

eess.AS2026

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challengi…

eess.AS2026

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang +7

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and l…

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +36

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

cs.CV2026

Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

Jaehyun Jang, Eunseop Yoon, Hee Suk Yoon +3

While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying solely on output token distributio…

cs.SD2026

Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?

Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson +2

Synthetic accented speech is a promising way to improve automatic speech recognition (ASR) when real accented recordings are scarce. We ask what makes such data useful for ASR fine…

eess.AS2026

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

SooHwan Eom, Hee Suk Yoon, Eunseop Yoon +2

Recent flow-matching text-to-speech (TTS) models, such as F5-TTS, rely on a reference transcript at inference time, obtained from an external ASR system. This dependency makes zero…