activity
20182023
most citedAn End-to-end Approach for Lexical Stress Detection based on Transformer

3 citations · 7 across the 9 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2023★ 2 cited

Audio Generation with Multiple Conditional Diffusion Model

Zhifang Guo, Jianguo Mao, Rui Tao +4

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To add…

cs.SD2023

Leveraging Language Model Capabilities for Sound Event Detection

Hualei Wang, Jianguo Mao, Zhifang Guo +3

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality,…

cs.SD2022

A Hybrid System of Sound Event Detection Transformer and Frame-wise Model for DCASE 2022 Task 4

Yiming Li, Zhifang Guo, Zhirong Ye +6

In this paper, we describe in detail our system for DCASE 2022 Task4. The system combines two considerably different models: an end-to-end Sound Event Detection Transformer (SEDT)…

cs.SD2020★ 2 cited

Learning generic feature representation with synthetic data for weakly-supervised sound event detection by inter-frame distance loss

Yuxin Huang, Liwei Lin, Xiangdong Wang +4

Due to the limitation of strong-labeled sound event detection data set, using synthetic data to improve the sound event detection system performance has been a new research focus.…

cs.SD2020

Guided multi-branch learning systems for sound event detection with sound separation

Yuxin Huang, Liwei Lin, Shuo Ma +5

In this paper, we describe in detail our systems for DCASE 2020 Task 4. The systems are based on the 1st-place system of DCASE 2019 Task 4, which adopts weakly-supervised framework…

cs.SD2019

Specialized Decision Surface and Disentangled Feature for Weakly-Supervised Polyphonic Sound Event Detection

Liwei Lin, Xiangdong Wang, Hong Liu +1

In this paper, a special decision surface for the weakly-supervised sound event detection (SED) and a disentangled feature (DF) for the multi-label problem in polyphonic SED are pr…