activity
20212026
most citedImproving Hybrid CTC/Attention End-to-end Speech Recognition with Pretrained Acoustic and Language Model

6 citations · 8 across the 14 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Sitong Cheng, Weizhen Bian, Songjun Cao +9

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue),…

cs.SD2024

FreeCodec: A disentangled neural speech codec with fewer tokens

Youqiang Zheng, Weiping Tu, Yueteng Kang +5

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…

cs.SD2024★ 1 cited

A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition

Yangze Li, Xiong Wang, Songjun Cao +3

Audio-LLM introduces audio modality into a large language model (LLM) to enable a powerful LLM to recognize, understand, and generate audio. However, during speech recognition in n…

cs.SD2023

Two Stage Contextual Word Filtering for Context bias in Unified Streaming and Non-streaming Transducer

Zhanheng Yang, Sining Sun, Xiong Wang +3

It is difficult for an E2E ASR system to recognize words such as entities appearing infrequently in the training data. A widely used method to mitigate this issue is feeding contex…

cs.SD2022

Censer: Curriculum Semi-supervised Learning for Speech Recognition Based on Self-supervised Pre-training

Bowen Zhang, Songjun Cao, Xiaoming Zhang +3

Recent studies have shown that the benefits provided by self-supervised pre-training and self-training (pseudo-labeling) are complementary. Semi-supervised fine-tuning strategies u…

cs.SD2022

Conversational Speech Recognition By Learning Conversation-level Characteristics

Kun Wei, Yike Zhang, Sining Sun +2

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can natura…