6 papers
A Generative-First Neural Audio Autoencoder
Jonah Casebeer, Ge Zhu, Zhepei Wang +1
Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single…
Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation
Ge Zhu, Yutong Wen, Zhiyao Duan
Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing…
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
Xin Wang, Héctor Delgado, Hemlata Tak +26
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce…
MusicHiFi: Fast High-Fidelity Stereo Vocoding
Ge Zhu, Juan-Pablo Caceres, Zhiyao Duan +1
Diffusion-based audio and music generation models commonly perform generation by constructing an image representation of audio (e.g., a mel-spectrogram) and then convert it to audi…
Cacophony: An Improved Contrastive Audio-Text Model
Ge Zhu, Jordan Darefsky, Zhiyao Duan
Despite recent advancements, audio-text models still lag behind their image-text counterparts in scale and performance. In this paper, we propose to improve both the data scale and…
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
Yinghao Aaron Li, Xilin Jiang, Jordan Darefsky +2
The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and c…