16 citations · 18 across the 6 of their papers we have counts for
6 papers
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition
Wenjing Zhu, Sining Sun, Changhao Shan +2
Conformer-based attention models have become the de facto backbone model for Automatic Speech Recognition tasks. A blank symbol is usually introduced to align the input and output…
Learning a Structural Causal Model for Intuition Reasoning in Conversation
Hang Chen, Bingyu Liao, Jing Luo +2
Reasoning, a crucial aspect of NLP research, has not been adequately addressed by prevailing models including Large Language Model. Conversation reasoning, as a critical component…
How to Enhance Causal Discrimination of Utterances: A Case on Affective Reasoning
Hang Chen, Jing Luo, Xinyu Yang +1
Our investigation into the Affective Reasoning in Conversation (ARC) task highlights the challenge of causal discrimination. Almost all existing models, including large language mo…
Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning
Zhaoxi Mu, Xinyu Yang, Wenjing Zhu
Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information o…
A Multi-Stage Triple-Path Method for Speech Separation in Noisy and Reverberant Environments
Zhaoxi Mu, Xinyu Yang, Xiangyuan Yang +1
In noisy and reverberant environments, the performance of deep learning-based speech separation methods drops dramatically because previous methods are not designed and optimized f…
Speech Emotion Recognition with Global-Aware Fusion on Multi-scale Feature Representation
Wenjing Zhu, Xiang Li
Speech Emotion Recognition (SER) is a fundamental task to predict the emotion label from speech data. Recent works mostly focus on using convolutional neural networks~(CNNs) to lea…