3 citations · 3 across the 1 of their papers we have counts for
4 papers · 1 filter
The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring
Ali Keramati, Mark Warschauer
Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat these in isolation, harming…
Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate
Ali Keramati, Justin Cheok, Jacob Horne +1
Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signa…
The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge
Ali Keramati, Justin Cheok, Jacob Horne +1
Multi-agent debate systems are typically evaluated only on whether the final answer is correct, overlooking the quality of the intermediate reasoning that debate is designed to pro…
Fantastic Questions and Where to Find Them: FairytaleQA -- An Authentic Dataset for Narrative Comprehension
Ying Xu, Dakuo Wang, Mo Yu +15
Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity o…