2 papers
cs.SD2024
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
Marco Comunità, Zhi Zhong, Akira Takahashi +7
Recent advances in generative models that iteratively synthesize audio clips sparked great success to text-to-audio synthesis (TTA), but with the cost of slow synthesis speed and h…
cs.SD2023
An Attention-based Approach to Hierarchical Multi-label Music Instrument Classification
Zhi Zhong, Masato Hirano, Kazuki Shimada +3
Although music is typically multi-label, many works have studied hierarchical music tagging with simplified settings such as single-label data. Moreover, there lacks a framework to…