papers

Publications (20)

eess.AS2022

Music Source Separation with Generative Flow

Ge Zhu, Jordan Darefsky, Fei Jiang +2

Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obta…

cs.SD2023

EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis

Ge Zhu, Yutong Wen, Marc-André Carbonneau +1

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. Thi…

eess.AS2025

ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech

Xin Wang, Héctor Delgado, Hemlata Tak +26

ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce…

cs.SD2024

Cacophony: An Improved Contrastive Audio-Text Model

Ge Zhu, Jordan Darefsky, Zhiyao Duan

Despite recent advancements, audio-text models still lag behind their image-text counterparts in scale and performance. In this paper, we propose to improve both the data scale and…

eess.AS2023

Transcription free filler word detection with Neural semi-CRFs

Ge Zhu, Yujia Yan, Juan-Pablo Caceres +1

Non-linguistic filler words, such as "uh" or "um", are prevalent in spontaneous speech and serve as indicators for expressing hesitation or uncertainty. Previous works for detectin…

cs.CV2026

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

Xinyue Xu, Zheng Zhang, Kunyang Ma +5

The paper introduces DM-KG, a direction‑metric knowledge graph that extracts 3D spatial relationships from street‑view images and injects them into vision‑language models to improv…

#spatial reasoning#vision-language models#street view imagery#knowledge graph
eess.AS2021

Y-Vector: Multiscale Waveform Encoder for Speaker Embedding

Ge Zhu, Fei Jiang, Zhiyao Duan

State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features. Recent studies attempted to extract speak…

cs.CV2024

Bridging the Gap: Aligning Text-to-Image Diffusion Models with Specific Feedback

Xuexiang Niu, Jinping Tang, Lei Wang +1

Learning from feedback has been shown to enhance the alignment between text prompts and images in text-to-image diffusion models. However, due to the lack of focus in feedback cont…

eess.AS2021

UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021

Xinhui Chen, You Zhang, Ge Zhu +1

In this paper, we present UR-AIR system submission to the logical access (LA) and the speech deepfake (DF) tracks of the ASVspoof 2021 Challenge. The LA and DF tasks focus on synth…

eess.AS2022

A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification

You Zhang, Ge Zhu, Zhiyao Duan

The performance of automatic speaker verification (ASV) systems could be degraded by voice spoofing attacks. Most existing works aimed to develop standalone spoofing countermeasure…

cs.SD2025

Presto! Distilling Steps and Layers for Accelerating Music Generation

Zachary Novack, Ge Zhu, Jonah Casebeer +3

Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration…

cs.CL2022

Filler Word Detection and Classification: A Dataset and Benchmark

Ge Zhu, Juan-Pablo Caceres, Justin Salamon

Filler words such as `uh' or `um' are sounds or words people use to signal they are pausing to think. Finding and removing filler words from recordings is a common and tedious task…

eess.AS2021

An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems

You Zhang, Ge Zhu, Fei Jiang +1

Spoofing countermeasure (CM) systems are critical in speaker verification; they aim to discern spoofing attacks from bona fide speech trials. In practice, however, acoustic conditi…

cs.CV2023

Sharp Eyes: A Salient Object Detector Working The Same Way as Human Visual Characteristics

Ge Zhu, Jinbao Li, Yahong Guo

Current methods aggregate multi-level features or introduce edge and skeleton to get more refined saliency maps. However, little attention is paid to how to obtain the complete sal…

cs.CL2024

Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation

Yinghao Aaron Li, Xilin Jiang, Jordan Darefsky +2

The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and c…

eess.AS2021

A study of the robustness of raw waveform based speaker embeddings under mismatched conditions

Ge Zhu, Frank Cwitkowitz, Zhiyao Duan

In this paper, we conduct a cross-dataset study on parametric and non-parametric raw-waveform based speaker embeddings through speaker verification experiments. In general, we obse…

cs.SD2024

MusicHiFi: Fast High-Fidelity Stereo Vocoding

Ge Zhu, Juan-Pablo Caceres, Zhiyao Duan +1

Diffusion-based audio and music generation models commonly perform generation by constructing an image representation of audio (e.g., a mel-spectrogram) and then convert it to audi…

cs.SD2026

Stemphonic: All-at-once Flexible Multi-stem Music Generation

Shih-Lun Wu, Ge Zhu, Juan-Pablo Caceres +2

Music stem generation, the task of producing musically-synchronized and isolated instrument audio clips, offers the potential of greater user control and better alignment with musi…

cs.SD2026

Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation

Ge Zhu, Yutong Wen, Zhiyao Duan

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing…

cs.SD2026

A Generative-First Neural Audio Autoencoder

Jonah Casebeer, Ge Zhu, Zhepei Wang +1

Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single…