5 papers · 1 filter
Training with Pseudo-Code for Instruction Following
Prince Kumar, Rudra Murthy, Riyaz Bhat +1
Despite rapid advances in the capabilities of Large Language Models (LLMs), they continue to struggle with following relatively simple and unambiguous instructions, particularly wh…
Reducing the Scope of Language Models
David Yunis, Siyu Huo, Chulaka Gunasekara +1
Large language models (LLMs) are deployed in a wide variety of user-facing applications. Typically, these deployments have some specific purpose, like answering questions grounded…
KCIF: Knowledge-Conditioned Instruction Following
Rudra Murthy, Praveen Venkateswaran, Prince Kumar +1
LLM evaluation benchmarks have traditionally separated the testing of knowledge/reasoning capabilities from instruction following. In this work, we study the interaction between kn…
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is…
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
Young-Suk Lee, Chulaka Gunasekara, Danish Contractor +2
We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxo…