2 papers
cs.HC2025
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
Yoonseo Choi, Eunhye Kim, Hyunwoo Kim +4
If 100 people issue the same search query, they may have 100 different goals. While existing work on user-centric AI evaluation highlights the importance of aligning systems with f…
cs.AI2025
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
Yoonsu Kim, Hyoungwook Jin, Hayeon Doh +6
Students' handwritten math work provides a rich resource for diagnosing cognitive skills, as it captures intermediate reasoning beyond final answers. We investigate how current lar…