2 papers
cs.AI2026
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +10
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs h…
cs.CV2026
AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education
Hanchen David Wang, Yilin Liu, Madison J. Lee +4
Assessing learner competency in clinical simulation requires expert observation that is time-intensive, difficult to scale, and subject to inter-rater variability. Vision-language…