18 papers
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Unggi Lee, Sookbun Lee, Yeil Jeong +3
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point soluti…
Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models
Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim +1
Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a…
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
Yeil Jeong, Youngjin Yoo, Jiyoung Bae +5
Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evaluation. We present \textit{Tea…
Reinforcement Learning for Special Education: Aligning LLM Tutors to Diverse Learners through Disability-Adaptive Training
Unggi Lee, Jihoi Na, Yeil Jeong +2
Large language models are increasingly deployed as intelligent tutors, yet research on aligning them for special education remains absent. Recent work has applied reinforcement lea…
The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
Shim Jaechang, Unggi Lee
Aligning large language models (LLMs) as math tutors typically demands costly reinforcement-learning (RL) training and external LLM judges. We ask whether a frozen model's internal…
BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization
Jeyeon Eo, Joo Young Kim, Ran Ju +2
BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental repositories that primarily…