3 papers
cs.AI2026
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Dhaval C. Patel, Kaoutar El Maghraoui, Shuxin Lin +58
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated d…
cs.CL2026
Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning
Yuval Shemla, Ayal Yakobe, Tanmay Agarwal +2
Large language models are increasingly used as planning components in agentic systems, but current tool-use pipelines often require full tool schemas to be included in every prompt…
cs.RO2025
ORB: Operating Room Bot, Automating Operating Room Logistics through Mobile Manipulation
Jinkai Qiu, Yungjun Kim, Gaurav Sethia +4
Efficiently delivering items to an ongoing surgery in a hospital operating room can be a matter of life or death. In modern hospital settings, delivery robots have successfully tra…