5 papers
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder +6
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce V…
Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
Praveen Venkateswaran, Danish Contractor
In many real-world applications, users rely on natural language instructions to guide large language models (LLMs) across a wide range of tasks. These instructions are often comple…
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
Rahul Atul Bhope, K. R. Jayaram, Praveen Venkateswaran +1
Federated Learning (FL) enables collaborative model training across decentralized clients without sharing raw data, yet faces significant challenges in real-world settings where cl…
KCIF: Knowledge-Conditioned Instruction Following
Rudra Murthy, Praveen Venkateswaran, Prince Kumar +1
LLM evaluation benchmarks have traditionally separated the testing of knowledge/reasoning capabilities from instruction following. In this work, we study the interaction between kn…
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
Rahul Atul Bhope, Praveen Venkateswaran, K. R. Jayaram +3
Developers using LLMs and LLM-based agents in their applications have provided plenty of anecdotal evidence that in-context-learning (ICL) is fragile. In this paper, we show that i…