2 citations · 3 across the 4 of their papers we have counts for
4 papers
Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics
Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le +2
Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, org…
WildClaims: Information Access Conversations in the Wild(Chat)
Hideaki Joko, Shakiba Amirshahi, Charles L. A. Clarke +1
The rapid advancement of Large Language Models (LLMs) has transformed conversational systems into practical tools used by millions. However, the nature and necessity of information…
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
Shakiba Amirshahi, Amin Bigdeli, Charles L. A. Clarke +1
Retrieval augmented generation (RAG) systems provide a method for factually grounding the responses of a Large Language Model (LLM) by providing retrieved evidence, or context, as…
AInsight: Augmenting Expert Decision-Making with On-the-Fly Insights Grounded in Historical Data
Mohammad Abolnejadian, Shakiba Amirshahi, Matthew Brehmer +1
In decision-making conversations, experts must navigate complex choices and make on-the-spot decisions while engaged in conversation. Although extensive historical data often exist…