6 citations · 11 across the 23 of their papers we have counts for
11 papers · 1 filter
Cached LLM Probability Retrieval for Speech Recognition
Sheng Li, Takahiro Shinozaki, Tatsuya Kawahara
Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating eve…
Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment
Sheng Li, Takahiro Shinozaki
Modern neural speech systems can generate intelligible waveforms, but they usually hide the physical speech-production state that produced the sound. Conversely, biomechanical voca…
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
Pengcheng Wang, Sheng Li, Takahiro Shinozaki
In this paper, we propose RAG-Boost (ST-ShinozakiLab Task I system), which enhances the baseline LLM-based ASR system of the MLC-SLM Challenge (task I) with a retrieval-augmented g…
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
Chao Tan, Sheng Li, Yang Cao +2
Federated Learning (FL) is a privacy-preserving approach that allows servers to aggregate distributed models transmitted from local clients rather than training on user data. More…
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
Yuka Ko, Sheng Li, Chao-Han Huck Yang +1
With the strong representational power of large language models (LLMs), generative error correction (GER) for automatic speech recognition (ASR) aims to provide semantic and phonet…
MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
Wangjin Zhou, Zhengdong Yang, Chenhui Chu +4
Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio…