3 papers
cs.LG2026
SCATR: Simple Calibrated Test-Time Ranking
Divya Shyamal, Marta KneževiÄ, Lan Tran +3
Test-time scaling (TTS) improves large language models (LLMs) by allocating additional compute at inference time. In practice, TTS is often achieved through parallel scaling: gener…
cs.CR2026
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
Jin Zhao, Marta KneževiÄ, Tanja Käser
Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work evaluates pedagogical quality…
cs.AI2026
REFINE: Real-world Exploration of Interactive Feedback and Student Behaviour
Fares Fawzi, Seyed Parsa Neshaei, Marta Knezevic +2
Formative feedback is central to effective learning, yet providing timely, individualised feedback at scale remains a persistent challenge. While recent work has explored the use o…