10 papers
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning
Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan +4
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challengi…
Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement
Xulin Fan, Juan Azcarreta, Ashutosh Pandey +5
Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on…
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +36
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations
Haolong Zheng, Siyin Wang, Xulin Fan +2
Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. However, most auditory large langu…
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
Xulin Fan, Vishal Sunder, Samuel Thomas +3
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer…
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Xilin Jiang, Qiaolin Wang, Junkai Wu +30
Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand suc…