collaborators

7 papers

cs.CL2026

Pretrained self-supervised speech models can recognize unseen consonants

Chihiro Taguchi, Éric Le Ferrand, Hirosi Nakagawa +4

Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their tra…

cs.CL2026

Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan

Chihiro Taguchi, Yukinori Takubo, David Chiang

Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among thes…

cs.CL2026

Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs

Chihiro Taguchi, Richard Sproat

We present a system that uses LLMs as a tool in the development of Constructed Languages -- ConLangs, which we call IASC (Interactive Agentic System for ConLangs). The system is mo…

cs.CL2025

Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-

Chihiro Taguchi, Seiji Maekawa, Nikita Bhutani

Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in open-domain question answering (QA). However, optimal exte…

cs.CL2025

Building Tailored Speech Recognizers for Japanese Speaking Assessment

Yotaro Kubo, Richard Sproat, Chihiro Taguchi +1

This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels…

cs.CL2025

Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark

Chihiro Taguchi, Seng Mai, Keita Kurabe +4

Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering…