15 citations · 16 across the 6 of their papers we have counts for
6 papers
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
Fan Qian, Jiqing Han, Jianchen Li +3
The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation…
Contrastive Loss Based Frame-wise Feature disentanglement for Polyphonic Sound Event Detection
Yadong Guan, Jiqing Han, Hongwei Song +4
Overlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critica…
A Glance is Enough: Extract Target Sentence By Looking at A keyword
Ying Shi, Dong Wang, Lantian Li +1
This paper investigates the possibility of extracting a target sentence from multi-talker speech using only a keyword as input. For example, in social security applications, the ke…
Time-weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
Jian Guan, Youde Liu, Qiaoxi Zhu +3
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can…
Using Auxiliary Tasks In Multimodal Fusion Of Wav2vec 2.0 And BERT For Multimodal Emotion Recognition
Dekai Sun, Yancheng He, Jiqing Han
The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pretrained models as…
FurcaNet: An end-to-end deep gated convolutional, long short-term memory, deep neural networks for single channel speech separation
Ziqiang Shi, Huibin Lin, Liu Liu +4
Deep gated convolutional networks have been proved to be very effective in single channel speech separation. However current state-of-the-art framework often considers training the…