1 citations · 2 across the 2 of their papers we have counts for
11 papers
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
Shrestha Mohanty, Negar Arabzadeh, Andrea Tupini +5
Seamless interaction between AI agents and humans using natural language remains a key goal in AI research. This paper addresses the challenges of developing interactive agents cap…
Ranked List Truncation for Large Language Model-based Re-Ranking
Chuan Meng, Negar Arabzadeh, Arian Askari +2
We study ranked list truncation (RLT) from a novel "retrieve-then-re-rank" perspective, where we optimize re-ranking by truncating the retrieved list (i.e., trim re-ranking candida…
A Comparison of Methods for Evaluating Generative IR
Negar Arabzadeh, Charles L. A. Clarke
Information retrieval systems increasingly incorporate generative components. For example, in a retrieval augmented generation (RAG) system, a retrieval component might provide a s…
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
Negar Arabzadeh, Julia Kiseleva, Qingyun Wu +5
The rapid development in the field of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents to assist humans in their…
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels
Negar Arabzadeh, Charles L. A. Clarke
The rapid advancement of natural language processing, information retrieval (IR), computer vision, and other technologies has presented significant challenges in evaluating the per…
Adapting Standard Retrieval Benchmarks to Evaluate Generated Answers
Negar Arabzadeh, Amin Bigdeli, Charles L. A. Clarke
Large language models can now directly generate answers to many factual questions without referencing external sources. Unfortunately, relatively little attention has been paid to…