Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
Quality-aware Masked Diffusion Transformer for Enhanced Music Generation
Chang Li, Ruoyu Wang, Lijuan Liu +7
Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this pr…
cs.SD2024
Multitask frame-level learning for few-shot sound event detection
Liang Zou, Genwei Yan, Ruoyu Wang +4
This paper focuses on few-shot Sound Event Detection (SED), which aims to automatically recognize and classify sound events with limited samples. However, prevailing methods method…