activity
20242026
most citedSLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

3 citations · 4 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

19 papers · 1 filter

eess.AS2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang +36

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…

eess.AS20261 cited

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

Xiaoyu Yang, Yifan Yang, Zengrui Jin +5

Self-supervised learning (SSL) has significantly advanced acoustic representation learning. However, most existing models are optimised for either speech or audio event understandi…

eess.AS2026

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning

Hui Wang, Yifan Yang, Zeyue Tian +8

Audio generation and audio-to-text understanding remain largely separate, with diffusion models dominating high-fidelity synthesis and autoregressive (AR) language models driving c…

eess.AS2026

Position: Towards Responsible Evaluation for Text-to-Speech

Yifan Yang, Hui Wang, Bing Han +4

Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, co…

eess.AS2026

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

Guanrou Yang, Tian Tan, Qian Chen +12

Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks current…

eess.AS2026

Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training

Yifan Yang, Bing Han, Hui Wang +8

Modeling fine-grained speaking styles remains challenging for language-speech representation pre-training, as existing speech-text models are typically trained with coarse captions…