2 citations · 4 across the 3 of their papers we have counts for
5 papers
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
Benjamin Elder, Anupama Murthi, Jungkoo Kang +4
Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in reali…
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
Daniel McDuff, Tim Korjakow, Kevin Klyman +1
Foundation models have had a transformative impact on AI. A combination of large investments in research and development, growing sources of digital data for training, and architec…
Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
Praveen Venkateswaran, Danish Contractor
In many real-world applications, users rely on natural language instructions to guide large language models (LLMs) across a wide range of tasks. These instructions are often comple…
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is…
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
Young-Suk Lee, Chulaka Gunasekara, Danish Contractor +2
We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxo…