2 citations · 3 across the 5 of their papers we have counts for
8 papers · 1 filter
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer
Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing a…
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
Asaf Yehudai, Lilach Eden, Yotam Perlitz +2
The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking,…
JuStRank: Benchmarking LLM Judges for System Ranking
Ariel Gera, Odellia Boni, Yotam Perlitz +3
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and ver…
CHAMP: Efficient Annotation and Consolidation of Cluster Hierarchies
Arie Cattan, Tom Hope, Doug Downey +4
Various NLP tasks require a complex hierarchical structure over nodes, where each node is a cluster of items. Examples include generating entailment graphs, hierarchical cross-docu…
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization
Arie Cattan, Lilach Eden, Yoav Kantor +1
Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments. KPA extracts the main points in the data as a list of c…
Every Bite Is an Experience: Key Point Analysis of Business Reviews
Roy Bar-Haim, Lilach Eden, Yoav Kantor +2
Previous work on review summarization focused on measuring the sentiment toward the main aspects of the reviewed product or business, or on creating a textual summary. These approa…