2 papers
cs.AI2026
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
Grégoire Martinon, Ibrahim Merad, Mohammed Raki
Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotation and biased LLM-as-judge p…
cs.AI2025
Towards a rigorous evaluation of RAG systems: the challenge of due diligence
Grégoire Martinon, Alexandra Lorenzo de Brionne, Jérôme Bohard +3
The rise of generative AI, has driven significant advancements in high-risk sectors like healthcare and finance. The Retrieval-Augmented Generation (RAG) architecture, combining la…