4 papers
Building Tailored Speech Recognizers for Japanese Speaking Assessment
Yotaro Kubo, Richard Sproat, Chihiro Taguchi +1
This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels…
Languages Still Left Behind: Toward a Better Multilingual Machine Translation Benchmark
Chihiro Taguchi, Seng Mai, Keita Kurabe +4
Multilingual machine translation (MT) benchmarks play a central role in evaluating the capabilities of modern MT systems. Among them, the FLORES+ benchmark is widely used, offering…
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-
Chihiro Taguchi, Seiji Maekawa, Nikita Bhutani
Retrieval-augmented generation (RAG) and long-context language models (LCLMs) both address context limitations of LLMs in open-domain question answering (QA). However, optimal exte…
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
Hiroyuki Deguchi, Go Kamoda, Yusuke Matsushita +4
Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that p…