activity
20242026
most citedSPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

eess.AS2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu +10

Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder…

eess.AS20261 cited

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

Xiaoyu Yang, Yifan Yang, Zengrui Jin +5

Self-supervised learning (SSL) has significantly advanced acoustic representation learning. However, most existing models are optimised for either speech or audio event understandi…

cs.HC2025

PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback

Sirui Chen, Jinsong Zhou, Xinli Xu +3

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized…

cs.CL2025

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Wenyi Yu, Siyin Wang, Xiaoyu Yang +7

In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such…

cs.CL2025

LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation

Haitao Li, Yifan Chen, Yiran Hu +7

Retrieval-augmented generation (RAG) has proven highly effective in improving large language models (LLMs) across various domains. However, there is no benchmark specifically desig…

eess.AS2025

MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events

Xiaoyu Yang, Qiujia Li, Chao Zhang +1

With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still chall…