most citedExploring Self-Supervised Contrastive Learning of Spatial Sound Event Representation

1 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20241 cited

Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation

Yinghao Aaron Li, Xilin Jiang, Jordan Darefsky +2

The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and c…

eess.AS2024

Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis

Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea +2

It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in mu…

eess.AS20231 cited

Exploring Self-Supervised Contrastive Learning of Spatial Sound Event Representation

Xilin Jiang, Cong Han, Yinghao Aaron Li +1

In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and…

eess.AS20231 cited

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

Yinghao Aaron Li, Cong Han, Xilin Jiang +1

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these network…

cs.CL2023

Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions

Yinghao Aaron Li, Cong Han, Xilin Jiang +1

Large-scale pre-trained language models have been shown to be helpful in improving the naturalness of text-to-speech (TTS) models by enabling them to produce more naturalistic pros…