activity
20242026
collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challengi…

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +36

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

eess.AS2026

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations

Haolong Zheng, Siyin Wang, Xulin Fan +2

Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large langu…

eess.AS2026

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions

Xulin Fan, Vishal Sunder, Samuel Thomas +3

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer…

eess.AS2025

Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities

Jialu Li, Marvin Lavechin, Xulin Fan +4

Naturalistic recordings capture audio in real-world environments where participants behave naturally without interference from researchers or experimental protocols. Naturalistic l…

eess.AS2025

ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

Zhongweiyang Xu, Xulin Fan, Zhong-Qiu Wang +2

Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse…