activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Pretrained self-supervised speech models can recognize unseen consonants

Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa +4

Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their tra…

cs.CL2026

Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan

Chihiro Taguchi, Yukinori Takubo, David Chiang

Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among thes…

cs.CL2025

Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark

Chihiro Taguchi, Seng Mai, Keita Kurabe +4

Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering…

cs.CL2024

Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't

Chihiro Taguchi, David Chiang

We investigate what linguistic factors affect the performance of Automatic Speech Recognition (ASR) models. We hypothesize that orthographic and phonological complexities both degr…

cs.CL2024

Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information

Chihiro Taguchi, Jefferson Saransig, Dayana Velásquez +1

This paper presents Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador. Kichwa is an extremely low-resource…