2 papers
cs.AI2026
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
Yoonsu Kim, Hyoungwook Jin, Hayeon Doh +6
Students' handwritten math work provides a rich resource for diagnosing cognitive skills, as it captures intermediate reasoning beyond final answers. We investigate how current lar…
cs.HC2025
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
Yoonseo Choi, Eunhye Kim, Hyunwoo Kim +4
If 100 people issue the same search query, they may have 100 different goals. While existing work on user-centric AI evaluation highlights the importance of aligning systems with f…