activity
20202026
most citedGeneral Agent Evaluation

2 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer

Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing a…

cs.CL2025

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

Asaf Yehudai, Lilach Eden, Yotam Perlitz +2

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking,…

cs.CL2024

JuStRank: Benchmarking LLM Judges for System Ranking

Ariel Gera, Odellia Boni, Yotam Perlitz +3

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and ver…

cs.CL2023

CHAMP: Efficient Annotation and Consolidation of Cluster Hierarchies

Arie Cattan, Tom Hope, Doug Downey +4

Various NLP tasks require a complex hierarchical structure over nodes, where each node is a cluster of items. Examples include generating entailment graphs, hierarchical cross-docu…

cs.CL20231 cited

From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization

Arie Cattan, Lilach Eden, Yoav Kantor +1

Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments. KPA extracts the main points in the data as a list of c…

cs.CL2021

Every Bite Is an Experience: Key Point Analysis of Business Reviews

Roy Bar-Haim, Lilach Eden, Yoav Kantor +2

Previous work on review summarization focused on measuring the sentiment toward the main aspects of the reviewed product or business, or on creating a textual summary. These approa…