6 papers
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
Adam Nohejl, Xuanxin Wu, Yusuke Ide +3
We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the open track, and an explainab…
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
Yusuke Ide, Adam Nohejl, Joshua Tanner +3
We study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Dictionary definitions are an essential resource for le…
CoAM: Corpus of All-Type Multiword Expressions
Yusuke Ide, Joshua Tanner, Adam Nohejl +4
Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machi…
Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries
Haruki Sakajo, Yusuke Ide, Justin Vasselli +4
Cross-lingual vocabulary transfer plays a promising role in adapting pre-trained language models to new languages, including low-resource languages. Existing approaches that utiliz…
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
Yusuke Ide, Yuto Nishida, Justin Vasselli +4
The grammatical knowledge of language models (LMs) is often measured using a benchmark of linguistic minimal pairs, where the LMs are presented with a pair of acceptable and unacce…
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
Kazuki Hayashi, Kazuma Onishi, Toma Suzuki +7
Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these mod…