29 papers
Relation Geometry in Semantic Space of Language Models
Zhihan Cao, Hiroaki Yamada, Simone Teufel +4
The paper investigates how different semantic relations are reflected in the geometric structure of word embedding spaces produced by various language models, examining region clus…
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
Kaito Watanabe, Taisei Yamamoto, Tomoki Doi +1
One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we…
YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese
Ryota Mibayashi, Hiroya Takamura, Hitomi Yanaka
We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character o…
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
Yuheng Lu, Qingcheng Zeng, Heli Qi +6
Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks,…
Revisiting the Systematicity in Negation in the Era of In-Context Learning
Hitomi Yanaka, Taisei Yamamoto
Understanding the meaning of negated sentences remains one of the challenges for language models, even in the era of large language models (LLMs). We analyze systematicity regardin…
Analysis of the Neglect-Zero Effect in Large Language Models
Jin Tanaka, Daiki Matsuoka, Ryoma Kumon +1
We investigate the extent to which the language processing of LLMs resembles human cognitive processes, focusing on a human cognitive bias called the …