7 papers
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
Natsuo Yamashita, Koichi Nagatsuka, Hiroaki Kokubo +2
End-to-end automatic speech recognition often degrades on domain-specific data due to scarce in-domain resources. We propose a synthetic-data-based domain adaptation framework with…
Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
Phurich Saengthong, Tomoya Nishida, Kota Dohi +2
Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems ex…
Can VLM Pseudo-Labels Train a Time-Series QA Model That Outperforms the VLM?
Takuya Fujimura, Kota Dohi, Natsuo Yamashita +1
Time-series question answering (TSQA) tasks face significant challenges due to the lack of labeled data. Alternatively, with recent advancements in large-scale models, vision-langu…
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
Natsuo Yamashita, Masaaki Yamamoto, Hiroaki Kokubo +1
Generative error correction (GER) with large language models (LLMs) has emerged as an effective post-processing approach to improve automatic speech recognition (ASR) performance.…
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
Natsuo Yamashita, Masaaki Yamamoto, Yohei Kawaguchi
Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especi…
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
Ryotaro Nagase, Takashi Sumiyoshi, Natsuo Yamashita +2
This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited t…