5 papers
Quantifying Retriever-Generator Alignment in RAG with Local Explanations
Korbinian Randl, Guido Rocchietti, Aron Henriksson +3
Retrieval-Augmented Generation (RAG) systems combine dense retrievers and language models to ground their outputs in external documents. However, the interaction between these comp…
Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems
Maria Movin, Claudia Hauff, Aron Henriksson +1
LLM-driven GUI agents are increasingly used in production systems to automate workflows and simulate users for evaluation and optimization. Yet most GUI-agent evaluations emphasize…
Efficient Text Classification with Conformal In-Context Learning
Ippokratis Pantelidis, Korbinian Randl, Aron Henriksson
Large Language Models (LLMs) demonstrate strong in-context learning abilities, yet their effectiveness in text classification depends heavily on prompt design and incurs substantia…
SemEval-2025 Task 9: The Food Hazard Detection Challenge
Korbinian Randl, John Pavlopoulos, Aron Henriksson +2
In this challenge, we explored text-based food hazard prediction with long tail distributed classes. The task was divided into two subtasks: (1) predicting whether a web text impli…
Evaluating the Reliability of Self-Explanations in Large Language Models
Korbinian Randl, John Pavlopoulos, Aron Henriksson +1
This paper investigates the reliability of explanations generated by large language models (LLMs) when prompted to explain their previous output. We evaluate two kinds of such self…