4 papers
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Unggi Lee, Sookbun Lee, Yeil Jeong +3
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point soluti…
An Auditable AI Agent Loop for Empirical Economics: A Case Study in Forecast Combination
Minchul Shin
AI coding agents, general purpose assistants that write and execute code, make empirical specification search fast and cheap, but they also widen hidden researcher degrees of freed…
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
Unggi Lee, Minchul Shin, Yeil Jeong +5
Aligning LLMs for math tutoring typically requires RL-based training with multi-GPU infrastructure. We investigate whether training-free prompt optimization-evolving only the syste…
At-Risk Transformation for U.S. Recession Prediction
Rahul Billakanti, Minchul Shin
We propose a simple binarization of predictors, an "at-risk" transformation, as an alternative to the standard practice of using continuous, standardized variables in recession for…