collaborators

8 papers

cs.CL2026

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

Sangwhan Moon, Daisuke Oba, Youmi Ma +2

Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare or unseen characters. We investi…

cs.CL2026

Neuron Level Analysis of Large Language Model in Legal Domain Reasoning

Eri Onami, Youmi Ma, Shuhei Kurita +1

We presented a neuron-level analysis of legal-domain reasoning in LLMs, comparing it with other applied domain tasks across seven open-weight models. Using neuron attribution score…

cs.CL2026

From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models

Youmi Ma, Naoaki Okazaki

Advances in mechanistic interpretability have identified special attention heads, known as retrieval heads, that are responsible for retrieving information from the context. Howeve…

cs.CL2026

Synthesizing Instruction-Tuning Datasets with Contrastive Decoding

Tatsuya Ichinose, Youmi Ma, Masanari Oi +2

Using responses generated by high-performing large language models (LLMs) for instruction tuning has become a widely adopted approach. However, the existing literature overlooks a…

cs.LG2026

Rewriting Pre-Training Data Boosts LLM Performance in Math and Code

Kazuki Fujii, Yukito Tajima, Sakae Mizuki +14

The performance of large language models (LLMs) in program synthesis and mathematical reasoning is fundamentally limited by the quality of their pre-training corpora. We introduce…

cs.AI2025

QuantumBench: A Benchmark for Quantum Problem Solving

Shunya Minami, Tatsuya Ishigaki, Ikko Hamamura +6

Large language models are now integrated into many scientific workflows, accelerating data analysis, hypothesis generation, and design space exploration. In parallel with this grow…