1 citations · 1 across the 3 of their papers we have counts for
5 papers
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +36
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
Yi Su, Jisheng Bai, Qisheng Xu +2
Audio-Language Models (ALMs), trained on paired audio-text data, are designed to process, understand, and reason about audio-centric multimodal content. Unlike traditional supervis…
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
Kele Xu, Yulu Fang, Boda Zhou +6
This paper examines audio self-supervised learning (SSL) through the alignment between pretraining objectives, architectural inductive biases, and downstream applications. Rather t…
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
Hui Geng, Yi Su, Han Yin +7
Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To adapt these models to downstream tasks w…
Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review
Kele Xu, Yifan Wang, Ming Feng +5
Human-computer interaction has traditionally relied on the acoustic channel, a dependency that introduces systemic vulnerabilities to environmental noise, privacy constraints, and…