4 papers
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
Adam Nohejl, Xuanxin Wu, Yusuke Ide +3
We describe two types of models for vocabulary difficulty prediction: a high-accuracy black-box model, which achieved the top shared task result in the open track, and an explainab…
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
Hui Huang, Xuanxin Wu, Muyun Yang +1
This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields fou…
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
Xuanxin Wu, Yuki Arase, Masaaki Nagata
Sentence simplification aims to modify a sentence to make it easier to read and understand while preserving the meaning. Different applications require distinct simplification poli…
An In-depth Evaluation of Large Language Models in Sentence Simplification with Error-based Human Assessment
Xuanxin Wu, Yuki Arase
Recent studies have used both automatic metrics and human evaluations to assess the simplification abilities of LLMs. However, the suitability of existing evaluation methodologies…