most citedAdvanced Long-Content Speech Recognition With Factorized Neural Transducer

8 citations · 8 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS2024

TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers

Yakun Song, Zhuo Chen, Xiaofei Wang +3

Neural codec language model (LM) has demonstrated strong capability in zero-shot text-to-speech (TTS) synthesis. However, the codec LM often suffers from limitations in inference s…

cs.SD2024

AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection

Anbai Jiang, Bing Han, Zhiqiang Lv +6

Large pre-trained models have demonstrated dominant performances in multiple areas, where the consistency between pre-training and fine-tuning is the key to success. However, few w…

cs.SD2024

The Interspeech 2024 Challenge on Speech Processing Using Discrete Units

Xuankai Chang, Jiatong Shi, Jinchuan Tian +7

Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors. Numerous studies have highlighted the e…

cs.SD2024

StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

Sen Liu, Yiwei Guo, Xie Chen +1

While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of arti…

cs.SD20248 cited

Advanced Long-Content Speech Recognition With Factorized Neural Transducer

Xun Gong, Yu Wu, Jinyu Li +4

In this paper, we propose two novel approaches, which integrate long-content information into the factorized neural transducer (FNT) based architecture in both non-streaming (refer…

cs.CL2023

Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning

Guanrou Yang, Ziyang Ma, Zhisheng Zheng +3

Recent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and…