1 citations · 1 across the 6 of their papers we have counts for
Showing 2026 · cs.CLShow all
3 papers · 2 filters
cs.CL2026
FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity
Tyrone White, Yuki Arase
Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones.…
cs.CL2026
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
Hui Huang, Xuanxin Wu, Muyun Yang +1
This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields fou…
cs.CL2026
AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification
Hui Huang, Muyun Yang, Yuki Arase
Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verificatio…