4 papers
F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
Manvi Agarwal, Changhong Wang, Gael Richard
While music remains a challenging domain for generative models like Transformers, recent progress has been made by exploiting suitable musically-informed priors. One technique to l…
A Hybrid Model for Weakly-Supervised Speech Dereverberation
Louis Bahrman, Mathieu Fontaine, Gael Richard
This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms re…
Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects
Victor Deng, Changhong Wang, Gael Richard +1
In recent years, foundation models have significantly advanced data-driven systems across various domains. Yet, their underlying properties, especially when functioning as feature…
Using Random Codebooks for Audio Neural AutoEncoders
Benoît Giniès, Xiaoyu Bie, Olivier Fercoq +1
Latent representation learning has been an active field of study for decades in numerous applications. Inspired among others by the tokenization from Natural Language Processing an…