5 papers · 1 filter
Grounded Decoding for Autoregressive Speech Enhancement via Adaptive Code-Space Grounding and Local LLM Refinement
Hao Shi, Yuan Gao, Zhaoheng Ni +4
Large language model (LLM)-based autoregressive speech enhancement (SE) produces natural speech using learned clean-speech priors, but may hallucinate content unsupported by the in…
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
Xugang Lu, Peng Shen, Yu Tsao +1
Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary…
Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition
Hao Shi, Yuan Gao, Xugang Lu +1
Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantia…
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR…
Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition
Peng Shen, Xugang Lu, Hisashi Kawai
Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper,…