4 papers
TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
Tommaso Bendinelli, Artur Dox, Christian Holz
LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems; however, their performance on these…
Exploring LLM Agents for Cleaning Tabular Machine Learning Datasets
Tommaso Bendinelli, Artur Dox, Christian Holz
High-quality, error-free datasets are a key ingredient in building reliable, accurate, and unbiased machine learning (ML) models. However, real world datasets often suffer from err…
GECOBench: A Gender-Controlled Text Dataset and Benchmark for Quantifying Biases in Explanations
Rick Wilming, Artur Dox, Hjalmar Schulz +3
Large pre-trained language models have become a crucial backbone for many downstream tasks in natural language processing (NLP), and while they are trained on a plethora of data co…
EXACT: Towards a platform for empirically benchmarking Machine Learning model explanation methods
Benedict Clark, Rick Wilming, Artur Dox +11
The evolving landscape of explainable artificial intelligence (XAI) aims to improve the interpretability of intricate machine learning (ML) models, yet faces challenges in formalis…