7 papers
Pretrained self-supervised speech models can recognize unseen consonants
Chihiro Taguchi, Ãric Le Ferrand, Hirosi Nakagawa +4
Modern pretrained self-supervised automatic speech recognition models are trained on large-scale audio data to encode speech into contextualized representations. However, their tra…
Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
Chihiro Taguchi, Yukinori Takubo, David Chiang
Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among thes…
Creating ConLangs to Probe the Metalinguistic Grammatical Knowledge of LLMs
Chihiro Taguchi, Richard Sproat
We present a system that uses LLMs as a tool in the development of Constructed Languages -- ConLangs, which we call IASC (Interactive Agentic System for ConLangs). The system is mo…
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-
Chihiro Taguchi, Seiji Maekawa, Nikita Bhutani
Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in open-domain question answering (QA). However, optimal exte…
Building Tailored Speech Recognizers for Japanese Speaking Assessment
Yotaro Kubo, Richard Sproat, Chihiro Taguchi +1
This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels…
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark
Chihiro Taguchi, Seng Mai, Keita Kurabe +4
Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering…