most citedHierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023

3 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD20241 cited

MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection

Pengfei Cai, Yan Song, Kang Li +2

Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still r…

cs.MM2023

Encoder-Decoder-Based Intra-Frame Block Partitioning Decision

Yucheng Jiang, Han Peng, Yan Song +3

The recursive intra-frame block partitioning decision process, a crucial component of the next-generation video coding standards, exerts significant influence over the encoding tim…

eess.AS20233 cited

Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023

Haotian Wang, Yuxuan Xi, Hang Chen +11

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as rob…

stat.ME20231 cited

Fast robust location and scatter estimation: a depth-based method

Maoyu Zhang, Yan Song, Wenlin Dai

The minimum covariance determinant (MCD) estimator is ubiquitous in multivariate analysis, the critical step of which is to select a subset of a given size with the lowest sample c…

eess.AS2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

Kang Li, Yan Song, Li-Rong Dai +3

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio t…