3 papers
cs.CL2026
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani +3
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling perform…
cs.CL2026
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
Yihan Hong, Huaiyuan Yao, Bolin Shen +3
Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challengin…
cs.AI2026
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
Huaiyuan Yao, Wanpeng Xu, Justin Turnau +2
Preparing high-quality instructional materials remains a labor-intensive process that often requires extensive coordination among teaching faculty, instructional designers, and tea…