most citedEvaluating Model Explanations without Ground Truth

5 citations · 5 across the 5 of their papers we have counts for

collaborators

6 papers

cs.AI2026

Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set

Kaivalya Rawal, Eoin Delaney, Zihao Fu +2

Explainable artificial intelligence (XAI) is concerned with producing explanations indicating the inner workings of models. For a Rashomon set of similarly performing models, expla…

cs.LG2025

FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models

Zihao Fu, Ryan Brown, Shun Shao +3

Text-to-image diffusion models, such as Stable Diffusion, have demonstrated remarkable capabilities in generating high-quality and diverse images from natural language prompts. How…

cs.LG2025

CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions

Zihao Fu, Ming Liao, Chris Russell +1

Large language models have achieved remarkable success but remain largely black boxes with poorly understood internal mechanisms. To address this limitation, many researchers have…

cs.LG2025

LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations

Harry Mayne, Ryan Othniel Kearns, Yushi Yang +4

To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated co…

cs.CR2025

Multi-use LLM Watermarking and the False Detection Problem

Zihao Fu, Chris Russell

Digital watermarking is a promising solution for mitigating some of the risks arising from the misuse of automatically generated text. These approaches either embed non-specific wa…

cs.AI20255 cited

Evaluating Model Explanations without Ground Truth

Kaivalya Rawal, Zihao Fu, Eoin Delaney +1

There can be many competing and contradictory explanations for a single model prediction, making it difficult to select which one to use. Current explanation evaluation frameworks…