3 citations · 4 across the 4 of their papers we have counts for
4 papers
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
Yinghao Aaron Li, Xilin Jiang, Jordan Darefsky +2
The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and c…
EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis
Ge Zhu, Yutong Wen, Marc-André Carbonneau +1
Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. Thi…
Transcription free filler word detection with Neural semi-CRFs
Ge Zhu, Yujia Yan, Juan-Pablo Caceres +1
Non-linguistic filler words, such as "uh" or "um", are prevalent in spontaneous speech and serve as indicators for expressing hesitation or uncertainty. Previous works for detectin…
Sharp Eyes: A Salient Object Detector Working The Same Way as Human Visual Characteristics
Ge Zhu, Jinbao Li, Yahong Guo
Current methods aggregate multi-level features or introduce edge and skeleton to get more refined saliency maps. However, little attention is paid to how to obtain the complete sal…