Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +19
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, a…
eess.AS2025
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha +3
This paper presents a new challenge that calls for zero-shot text-to-speech (TTS) systems to augment speech data for the downstream task, personalized speech enhancement (PSE), as…
eess.AS2025
Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation
Jackie Lin, Georg Götz, Hermes Sampedro Llopis +9
This paper describes the synthesis of the room acoustics challenge as a part of the generative data augmentation workshop at ICASSP 2025. The challenge defines a unique generative…