3 papers
cs.CL2026
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
Yo Ehara
Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires s…
cs.CL2026
Accurate and Efficient Statistical Testing for Word Semantic Breadth
Yo Ehara
Measuring the breadth of a word's meaning, or its spread across contexts, has become feasible with contextualized token embeddings. A word type can be represented as a cloud of tok…
cs.AI2025
Educational Cone Model in Embedding Vector Spaces
Yo Ehara
Human-annotated datasets with explicit difficulty ratings are essential in intelligent educational systems. Although embedding vector spaces are widely used to represent semantic c…