activity
20182026
most citedConstructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

5 citations · 15 across the 16 of their papers we have counts for

collaborators

23 papers

cs.CL2026

Large Language Models as Modal Models in Linguistics

Haruto Suzuki, Saku Sugawara

The rapid advancement of large language models (LLMs) has intensified debates about their significance for linguistic theory. These debates are commonly divided into three position…

cs.CL2026

A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models

Rei Emura, Saku Sugawara

Language models (LMs) behave more like humans when their cognitive resources are restricted, particularly in predicting sentence processing costs such as reading times. However, it…

cs.CL2026

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

Akira Kawabata, Saku Sugawara

Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. However, most existing method…

cs.CL2026

CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models

Miyu Oba, Saku Sugawara

Recent work has examined language models from a linguistic perspective to better understand how they acquire language. Most existing benchmarks focus on judging grammatical accepta…

cs.CL2025

TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?

Yiwei Liu, Emma Jane Pretty, Jiahao Huang +1

While recent studies explore Large Language Models' (LLMs) performance on Theory of Mind (ToM) reasoning tasks, research on ToM abilities that require more nuanced social context i…

cs.CL2025

Specification-Aware Machine Translation and Evaluation for Purpose Alignment

Yoko Kayano, Saku Sugawara

In professional settings, translation is guided by communicative goals and client needs, often formalized as specifications. While existing evaluation frameworks acknowledge the im…