5 papers
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3
The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and test…
Agent Mentor: Framing Agent Knowledge through Semantic Trajectory Analysis
Roi Ben-Gigi, Yuval David, Fabiana Fournier +4
AI agent development relies heavily on natural language prompting to define agents' tasks, knowledge, and goals. These prompts are interpreted by Large Language Models (LLMs), whic…
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
Hadar Mulian, Sergey Zeltyn, Ido Levy +3
We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes f…
Selecting the Right LLM for eGov Explanations
Lior Limonad, Fabiana Fournier, Hadar Mulian +3
The perceived quality of the explanations accompanying e-government services is key to gaining trust in these institutions, consequently amplifying further usage of these services.…
Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems
Dany Moshkovich, Hadar Mulian, Sergey Zeltyn +3
The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluat…