Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
Yo Ehara
Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires s…
cs.CL2026
Accurate and Efficient Statistical Testing for Word Semantic Breadth
Yo Ehara
Measuring the breadth of a word's meaning, or its spread across contexts, has become feasible with contextualized token embeddings. A word type can be represented as a cloud of tok…