papers

Publications (7)

cs.CL2026

AutoMixer: Checkpoint Artifacts as Automatic Data Mixers

Ernie Chang, Yang Li, Patrick Huber +4

In language model training, it is desirable to equip models with capabilities from various tasks. However, it is not clear how to directly obtain the right data mixtures for these…

cs.SD2024

MusicFlow: Cascaded Flow Matching for Text Guided Music Generation

K R Prajwal, Bowen Shi, Matthew Lee +8

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music aud…

eess.AS2023

Self-Supervised Representations for Singing Voice Conversion

Tejas Jayashankar, Jilong Wu, Leda Sari +3

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio r…

eess.AS2023

Stack-and-Delay: a new codebook pattern for music generation

Gael Le Lan, Varun Nagaraja, Ernie Chang +5

In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…

eess.AS2024

High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Gael Le Lan, Bowen Shi, Zhaoheng Ni +9

We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48…

cs.SD2024

Simple and Controllable Music Generation

Jade Copet, Felix Kreuk, Itai Gat +5

We tackle the task of conditional music generation. We introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representati…