activity
20122022
most citedFurcaNet: An end-to-end deep gated convolutional, long short-term memory, deep neural networks for single channel speech separation

15 citations · 60 across the 11 of their papers we have counts for

collaborators

14 papers

cs.SD20222 cited

Contrastive Regularization for Multimodal Emotion Recognition Using Audio and Text

Fan Qian, Jiqing Han

Speech emotion recognition is a challenge and an important step towards more natural human-computer interaction (HCI). The popular approach is multimodal emotion recognition based…

eess.AS2022

Exploring Transformer's potential on automatic piano transcription

Longshen Ou, Ziyi Guo, Emmanouil Benetos +2

Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic…

cs.SD2020

Can We Trust Deep Speech Prior?

Ying Shi, Haolin Chen, Zhiyuan Tang +3

Recently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) arc…

eess.AS20201 cited

Speech Separation Based on Multi-Stage Elaborated Dual-Path Deep BiLSTM with Auxiliary Identity Loss

Ziqiang Shi, Rujie Liu, Jiqing Han

Deep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation.…

cs.SD2020

LaFurca: Iterative Refined Speech Separation Based on Context-Aware Dual-Path Parallel Bi-LSTM

Ziqiang Shi, Rujie Liu, Jiqing Han

Deep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation,…

cs.SD2019

Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events

Hongwei Song, Jiqing Han, Shiwen Deng +1

In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from e…