4 papers
Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
Sanyuan Chen, Min-Jae Hwang, Sho Inoue +12
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer traine…
From Real to Cloned Singer Identification
Dorian Desblancs, Gabriel Meseguer-Brocal, Romain Hennequin +1
Cloned voices of popular singers sound increasingly realistic and have gained popularity over the past few years. They however pose a threat to the industry due to personality righ…
An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging
Gabriel Meseguer-Brocal, Dorian Desblancs, Romain Hennequin
Self-supervised learning has emerged as a powerful way to pre-train generalizable machine learning models on large amounts of unlabeled data. It is particularly compelling in the m…
Music Augmentation and Denoising For Peak-Based Audio Fingerprinting
Kamil Akesbi, Dorian Desblancs, Benjamin Martin
Audio fingerprinting is a well-established solution for song identification from short recording excerpts. Popular methods rely on the extraction of sparse representations, general…