MULTIMODAL ANALYSIS: Informed content estimation and audio source separation
arXiv:2104.13276
Abstract
This dissertation proposes the study of multimodal learning in the context of musical signals. Throughout, we focus on the interaction between audio signals and text information. Among the many text sources related to music that can be used (e.g. reviews, metadata, or social network feedback), we concentrate on lyrics. The singing voice directly connects the audio signal and the text information in a unique way, combining melody and lyrics where a linguistic dimension complements the abstraction of musical instruments. Our study focuses on the audio and lyrics interaction for targeting source separation and informed content estimation.
Ph.D. dissertation. Thesis supervisor: Geoffroy Peeters. Jury:Laurent Girin, Gaël Richard, Rachel Bittner, Elena Cabrio, Bruno Gas, Perfecto Herrera Boyer, Antoine Liutkus
References in corpus (11)
- WaveNet: A Generative Model for Raw Audio
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Learning to Reweight Examples for Robust Deep Learning
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- One Model To Learn Them All
- Learning Multi-modal Similarity
- Singing voice separation: a study on training data
- Content based singing voice source separation via strong conditioning using aligned phonemes
- Data Cleansing with Contrastive Learning for Vocal Note Event Annotations