3 citations · 5 across the 5 of their papers we have counts for
5 papers
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
Pengfei Cai, Yan Song, Kang Li +2
Sound event detection (SED) methods that leverage a large pre-trained Transformer encoder network have shown promising performance in recent DCASE challenges. However, they still r…
Encoder-Decoder-Based Intra-Frame Block Partitioning Decision
Yucheng Jiang, Han Peng, Yan Song +3
The recursive intra-frame block partitioning decision process, a crucial component of the next-generation video coding standards, exerts significant influence over the encoding tim…
Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
Haotian Wang, Yuxuan Xi, Hang Chen +11
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as rob…
Fast robust location and scatter estimation: a depth-based method
Maoyu Zhang, Yan Song, Wenlin Dai
The minimum covariance determinant (MCD) estimator is ubiquitous in multivariate analysis, the critical step of which is to select a subset of a given size with the lowest sample c…
AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
Kang Li, Yan Song, Li-Rong Dai +3
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio t…