activity
20242026
collaborators

13 papers

cs.SD2026

Frontend Token Enhancement for Token-Based Speech Recognition

Takanori Ashihara, Shota Horiguchi, Kohei Matsuura +2

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and sp…

eess.AS2025

MOVER: Combining Multiple Meeting Recognition Systems

Naoyuki Kamo, Tsubasa Ochiai, Marc Delcroix +1

In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to c…

eess.AS2025

Generic Speech Enhancement with Self-Supervised Representation Space Loss

Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix +3

Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, t…

eess.AS2025

Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge

Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…

cs.CL2025

TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

Junyi Peng, Takanori Ashihara, Marc Delcroix +4

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previ…

eess.AS2025

Microphone Array Signal Processing and Deep Learning for Speech Enhancement

Reinhold Haeb-Umbach, Tomohiro Nakatani, Marc Delcroix +2

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal…