15 citations · 22 across the 11 of their papers we have counts for
11 papers
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
Frank Seide, Morrie Doulaty, Yangyang Shi +3
We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming. This is the f…
FADI-AEC: Fast Score Based Diffusion Model Guided by Far-end Signal for Acoustic Echo Cancellation
Yang Liu, Li Wan, Yun Li +5
Despite the potential of diffusion models in speech enhancement, their deployment in Acoustic Echo Cancellation (AEC) has been restricted. In this paper, we propose DI-AEC, pioneer…
On The Open Prompt Challenge In Conditional Audio Generation
Ernie Chang, Sidd Srinivasan, Mahi Luthra +8
Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is ch…
In-Context Prompt Editing For Conditional Audio Generation
Ernie Chang, Pin-Jie Lin, Yang Li +6
Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-au…
FoleyGen: Visually-Guided Audio Generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4
Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) gen…
Stack-and-Delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang +5
In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner…