collaborators

5 papers

cs.CL2026

SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora

Masataka Yoneda, Yusuke Matsushita, Go Kamoda +4

We present SoftMatcha 2, an ultra-fast and flexible search algorithm that enables search over trillion-scale natural language corpora in under 0.3 seconds while allowing semantic v…

cs.CL2026

Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics

Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi +4

Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we study how LMs compare such quan…

cs.CL2025

Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki

Mutsumi Sasaki, Go Kamoda, Ryosuke Takahashi +4

Temporal reasoning and knowledge are essential capabilities for language models (LMs). While much prior work has analyzed and improved temporal reasoning in LMs, most studies have…

cs.CL2025

SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches

Hiroyuki Deguchi, Go Kamoda, Yusuke Matsushita +4

Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that p…

cs.CL2025

Weight-based Analysis of Detokenization in Language Models: Understanding the First Stage of Inference Without Inference

Go Kamoda, Benjamin Heinzerling, Tatsuro Inaba +3

According to the stages-of-inference hypothesis, early layers of language models map their subword-tokenized input, which does not necessarily correspond to a linguistically meanin…