10 papers
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
Isabel Molnar, Peiyu Li, Si Chen +5
Higher education instructors often lack timely and pedagogically grounded support, as scalable instructional guidance remains limited and existing tools rely on generic chatbot adv…
Building Scaffolding Dialogue Data with LLM-Simulated Novices
Si Chen, Izzy Molnar, Ting Hua +6
High-quality, multi-turn instructional dialogues between novices and experts are essential for developing AI systems that support teaching, learning, and decision-making. These dia…
A Human-Centred AI System for Multi-Actor Planning and Collaboration in Family Learning
Si Chen, Jingyi Xie, Yao Li +8
Family learning takes place in everyday routines where children and caregivers read, practice, and develop new skills together. Despite growing interest in AI tutors, most existing…
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
Si Chen, Le Huy Khiem, Annalisa Szymanski +3
Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domai…
Building AI Literacy at Home: How Families Navigate Children's Self-Directed Learning with AI
Jingyi Xie, Chuhao Wu, Ge Wang +4
As generative AI becomes embedded in children's learning spaces, families face new challenges in guiding its use. Middle childhood (ages 7-13) is a critical stage where children se…