activity
20212026
most citedPolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models

1 citations · 2 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2025

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

Runyan Yang, Yuke Si, Yingying Gao +3

While large audio language models excel at tasks like ASR and emotion recognition, they still struggle with complex reasoning due to the modality gap between audio and text as well…

eess.AS2025

HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling

Yuke Si, Runyan Yang, Yingying Gao +3

Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared archi…

eess.AS2024

GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model

Yingying Gao, Shilei Zhang, Chao Deng +1

Pre-trained speech language models such as HuBERT and WavLM leverage unlabeled speech data for self-supervised learning and offer powerful representations for numerous downstream t…

eess.AS2024

Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network

Yanan Chen, Zihao Cui, Yingying Gao +3

The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due…

eess.AS2023

GenDistiller: Distilling Pre-trained Language Models based on Generative Models

Yingying Gao, Shilei Zhang, Zihao Cui +3

Self-supervised pre-trained models such as HuBERT and WavLM leverage unlabeled speech data for representation learning and offer significantly improve for numerous downstream tasks…

eess.AS2022

Multiple Confidence Gates For Joint Training Of SE And ASR

Tianrui Wang, Weibin Zhu, Yingying Gao +2

Joint training of speech enhancement model (SE) and speech recognition model (ASR) is a common solution for robust ASR in noisy environments. SE focuses on improving the auditory q…