8 citations · 15 across the 15 of their papers we have counts for
Showing 2024 · cs.SDShow all
2 papers · 2 filters
cs.SD2024
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
K R Prajwal, Bowen Shi, Matthew Lee +8
We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music aud…
cs.SD2024
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
HyoJung Han, Mohamed Anwar, Juan Pino +4
Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potent…