3 papers
cs.SD2023
HCLAS-X: Hierarchical and Cascaded Lyrics Alignment System Using Multimodal Cross-Correlation
Minsung Kang, Soochul Park, Keunwoo Choi
In this work, we address the challenge of lyrics alignment, which involves aligning the lyrics and vocal components of songs. This problem requires the alignment of two distinct mo…
eess.AS2023
A Demand-Driven Perspective on Generative Audio AI
Sangshin Oh, Minsung Kang, Hyeongi Moon +2
To achieve successful deployment of AI research, it is crucial to understand the demands of the industry. In this paper, we present the results of a survey conducted with professio…
eess.AS2023
FALL-E: A Foley Sound Synthesis Model and Strategies
Minsung Kang, Sangshin Oh, Hyeongi Moon +2
This paper introduces FALL-E, a foley synthesis system and its training/inference strategies. The FALL-E model employs a cascaded approach comprising low-resolution spectrogram gen…