17 citations · 23 across the 13 of their papers we have counts for
11 papers · 1 filter
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
Sara Rosenthal, Yannis Katsis, Vraj Shah +3
We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augmented generation, a popular use of large language models. We release a benchmark of 666 t…
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is…
DELIFT: Data Efficient Language model Instruction Fine Tuning
Ishika Agarwal, Krishnateja Killamsetty, Lucian Popa +1
Fine-tuning large language models (LLMs) is essential for enhancing their performance on specific tasks but is often resource-intensive due to redundant or uninformative data. To a…
Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains
Yu Zhang, Yunyi Zhang, Yanzhen Shen +5
Accurately typing entity mentions from text segments is a fundamental task for various natural language processing applications. Many previous approaches rely on massive human-anno…
Long-form Question Answering: An Iterative Planning-Retrieval-Generation Approach
Pritom Saha Akash, Kashob Kumar Roy, Lucian Popa +1
Long-form question answering (LFQA) poses a challenge as it involves generating detailed answers in the form of paragraphs, which go beyond simple yes/no responses or short factual…
Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language Explanations
Bingsheng Yao, Prithviraj Sen, Lucian Popa +2
Human-annotated labels and explanations are critical for training explainable NLP models. However, unlike human-annotated labels whose quality is easier to calibrate (e.g., with a…