collaborators

8 papers

cs.AI2026

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder +6

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce V…

cs.CL2026

Training with Pseudo-Code for Instruction Following

Prince Kumar, Rudra Murthy, Riyaz Bhat +1

Despite rapid advances in the capabilities of Large Language Models (LLMs), they continue to struggle with following relatively simple and unambiguous instructions, particularly wh…

cs.LG2026

Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering

Praveen Venkateswaran, Danish Contractor

In many real-world applications, users rely on natural language instructions to guide large language models (LLMs) across a wide range of tasks. These instructions are often comple…

cs.SE2026

Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling

Benjamin Elder, Anupama Murthi, Jungkoo Kang +4

Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in reali…

cs.CL2025

Reducing the Scope of Language Models

David Yunis, Siyu Huo, Chulaka Gunasekara +1

Large language models (LLMs) are deployed in a wide variety of user-facing applications. Typically, these deployments have some specific purpose, like answering questions grounded…

cs.CY2025

New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses

Daniel McDuff, Tim Korjakow, Kevin Klyman +1

Foundation models have had a transformative impact on AI. A combination of large investments in research and development, growing sources of digital data for training, and architec…