3 papers
cs.AI2026
Holistic Evaluation and Failure Diagnosis of AI Agents
Netta Madvil, Gilad Dym, Alon Mecilati +12
AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches s…
cs.AI2026
Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)
Assaf Gerner, Netta Madvil, Nadav Barak +11
Large Language Models (LLMs) augmented with Retrieval-Augmented Generation (RAG) techniques are revolutionizing applications across multiple domains, such as healthcare, finance, a…
cs.LG2025
ORION Grounded in Context: Retrieval-Based Method for Hallucination Detection
Assaf Gerner, Netta Madvil, Nadav Barak +10
Despite advancements in grounded content generation, production Large Language Models (LLMs) based applications still suffer from hallucinated answers. We present "Grounded in Cont…