activity
20242026
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

Tyrone White, Yuki Arase

Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones.…

cs.CL2026

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases

Hui Huang, Xuanxin Wu, Muyun Yang +1

This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields fou…

cs.CL2026

AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification

Hui Huang, Muyun Yang, Yuki Arase

Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verificatio…

cs.CL2025

Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge

Xuanxin Wu, Yuki Arase, Masaaki Nagata

Sentence simplification aims to modify a sentence to make it easier to read and understand while preserving the meaning. Different applications require distinct simplification poli…

cs.CL2025

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Tarek Naous, Anagha Savit, Carlos Rafael Catalan +17

As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly important. Prior work by Naous et…

cs.CL2025

Aligning Sentence Simplification with ESL Learner's Proficiency for Language Acquisition

Guanlin Li, Yuki Arase, Noel Crespi

Text simplification is crucial for improving accessibility and comprehension for English as a Second Language (ESL) learners. This study goes a step further and aims to facilitate…