3 papers
cs.CL2025
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
Tuan Vu Ho, Hiroaki Kokubo, Masaaki Yamamoto +1
End-to-end automatic speech recognition (ASR) systems based on transformer architectures, such as Whisper, offer high transcription accuracy and robustness. However, their autoregr…
cs.SD2025
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
Natsuo Yamashita, Masaaki Yamamoto, Hiroaki Kokubo +1
Generative error correction (GER) with large language models (LLMs) has emerged as an effective post-processing approach to improve automatic speech recognition (ASR) performance.…
cs.SD2024
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
Natsuo Yamashita, Masaaki Yamamoto, Yohei Kawaguchi
Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especi…