3 citations · 4 across the 3 of their papers we have counts for
5 papers
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
Xiaochuan Li, Ke Wang, Girija Gouda +5
As Large Language Models (LLMs) become integrated into high-stakes domains, there is a growing need for evaluation methods that are both scalable for real-time deployment and relia…
Sequential Harmful Shift Detection Without Labels
Salim I. Amoukou, Tom Bewley, Saumitra Mishra +3
We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requi…
Interpretable LLM-based Table Question Answering
Giang Nguyen, Ivan Brugere, Shubham Sharma +3
Interpretability in Table Question Answering (Table QA) is critical, especially in high-stakes domains like finance and healthcare. While recent Table QA approaches based on Large…
Interpreting Language Reward Models via Contrastive Explanations
Junqi Jiang, Tom Bewley, Saumitra Mishra +2
Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM respons…
Privacy-Preserving Algorithmic Recourse
Sikha Pentyala, Shubham Sharma, Sanjay Kariyappa +2
When individuals are subject to adverse outcomes from machine learning models, providing a recourse path to help achieve a positive outcome is desirable. Recent work has shown that…