10 citations · 16 across the 9 of their papers we have counts for
4 papers · 1 filter
FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment
Betty Xiong, Jillian Fisher, Benjamin Newman +5
We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Adm…
Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
Danielle R. Thomas, Conrad Borchers, Jionghao Lin +6
Tutoring improves student achievement, but identifying and studying what tutoring actions are most associated with student learning at scale based on audio transcriptions is an ope…
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
Danielle R. Thomas, Conrad Borchers, Shambhavi Bhushan +3
Large language models (LLMs) are increasingly used to generate feedback, yet their impact on learning remains underexplored, especially compared to existing feedback methods. This…
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses
Jionghao Lin, Zifei Han, Danielle R. Thomas +4
One-on-one tutoring is widely acknowledged as an effective instructional method, conditioned on qualified tutors. However, the high demand for qualified tutors remains a challenge,…