activity
20242026
collaborators

9 papers

cs.CL2026

Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries

Yusuke Ide, Adam Nohejl, Joshua Tanner +3

We study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Dictionary definitions are an essential resource for le…

cs.CL2025

Multilingual Dialogue Generation and Localization with Dialogue Act Scripting

Justin Vasselli, Eunike Andriani Kardinata, Yusuke Sakai +1

Non-English dialogue datasets are scarce, and models are often trained or evaluated on translations of English-language dialogues, an approach which can introduce artifacts that re…

cs.CL2025

CoAM: Corpus of All-Type Multiword Expressions

Yusuke Ide, Joshua Tanner, Adam Nohejl +4

Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machi…

cs.CL2025

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…

cs.CL2025

Measuring the Robustness of Reference-Free Dialogue Evaluation Systems

Justin Vasselli, Adam Nohejl, Taro Watanabe

Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative respons…

cs.CL2025

Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity

Adam Nohejl, Taro Watanabe

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We eval…