2 citations · 3 across the 11 of their papers we have counts for
3 papers · 1 filter
Sequential Harmful Shift Detection Without Labels
Salim I. Amoukou, Tom Bewley, Saumitra Mishra +3
We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requi…
Interpreting Language Reward Models via Contrastive Explanations
Junqi Jiang, Tom Bewley, Saumitra Mishra +2
Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM respons…
Counterfactual Metarules for Local and Global Recourse
Tom Bewley, Salim I. Amoukou, Saumitra Mishra +2
We introduce T-CREx, a novel model-agnostic method for local and global counterfactual explanation (CE), which summarises recourse options for both individuals and groups in the fo…