3 papers
cs.AI2026
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
Hongliu Cao, Ilias Driouich, Eoin Thomas
Large Language Model (LLM)-based agents are increasingly adopted in high-stakes settings, but current benchmarks evaluate mainly whether a task was completed, not how. We introduce…
cs.CL2025
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
Ilias Driouich, Hongliu Cao, Eoin Thomas
Retrieval-augmented generation (RAG) systems improve large language model outputs by incorporating external knowledge, enabling more informed and context-aware responses. However,…
cs.CL2025
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
Hongliu Cao, Ilias Driouich, Robin Singh +1
Large Language Models (LLMs) have demonstrated impressive performance across diverse domains, yet they still encounter challenges such as insufficient domain-specific knowledge, bi…