From the 2 of 5 linked papers with an AI index.
5 papers
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
Peiyu Li, Xiaobao Huang, Ting Hua +1
The paper introduces CrochetBench, a benchmark that tests vision-language models on their ability to recognize crochet stitches, ground instructions, and generate executable croche…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Peiyu Li, Xiuxiu Tang, Si Chen +4
The paper proposes ATLAS, an adaptive testing framework using Item Response Theory to evaluate large language models more efficiently by selecting informative items, reducing requi…
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
Isabel Molnar, Peiyu Li, Si Chen +5
Higher education instructors often lack timely and pedagogically grounded support, as scalable instructional guidance remains limited and existing tools rely on generic chatbot adv…
Building Scaffolding Dialogue Data with LLM-Simulated Novices
Si Chen, Izzy Molnar, Ting Hua +6
High-quality, multi-turn instructional dialogues between novices and experts are essential for developing AI systems that support teaching, learning, and decision-making. These dia…
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
Si Chen, Isabel R. Molnar, Peiyu Li +5
Large language models (LLMs) typically generate direct answers, yet they are increasingly used as learning tools. Studying instructors' usage is critical, given their role in teach…