8 papers
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder +6
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce V…
Training with Pseudo-Code for Instruction Following
Prince Kumar, Rudra Murthy, Riyaz Bhat +1
Despite rapid advances in the capabilities of Large Language Models (LLMs), they continue to struggle with following relatively simple and unambiguous instructions, particularly wh…
Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
Praveen Venkateswaran, Danish Contractor
In many real-world applications, users rely on natural language instructions to guide large language models (LLMs) across a wide range of tasks. These instructions are often comple…
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
Benjamin Elder, Anupama Murthi, Jungkoo Kang +4
Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in reali…
Reducing the Scope of Language Models
David Yunis, Siyu Huo, Chulaka Gunasekara +1
Large language models (LLMs) are deployed in a wide variety of user-facing applications. Typically, these deployments have some specific purpose, like answering questions grounded…
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
Daniel McDuff, Tim Korjakow, Kevin Klyman +1
Foundation models have had a transformative impact on AI. A combination of large investments in research and development, growing sources of digital data for training, and architec…