14 papers
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
Gilad Nurko, Roi Benita, Yehoshua Dissen +4
Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separa…
Frontend Token Enhancement for Token-Based Speech Recognition
Takanori Ashihara, Shota Horiguchi, Kohei Matsuura +2
Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and sp…
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
Anselm Lohmann, Tomohiro Nakatani, Rintaro Ikeshita +3
Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially dis…
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
Shota Horiguchi, Naohiro Tawara, Takanori Ashihara +2
Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large dat…
MOVER: Combining Multiple Meeting Recognition Systems
Naoyuki Kamo, Tsubasa Ochiai, Marc Delcroix +1
In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to c…
Generic Speech Enhancement with Self-Supervised Representation Space Loss
Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix +3
Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, t…