◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Roy Bar-Haim

5 papers hereh-index 5905 citations13 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL4
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CLShow all

4 papers · 1 filter

cs.CL2026

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Noam Koren, Roy Bar-Haim, Abigail Goldsteen

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsi…

cs.CL2025

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

Noy Sternlicht, Ariel Gera, Roy Bar-Haim +2

We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multi…

cs.CL2025

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

Asaf Yehudai, Lilach Eden, Yotam Perlitz +2

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking,…

cs.CL2025

JuStRank: Benchmarking LLM Judges for System Ranking

Ariel Gera, Odellia Boni, Yotam Perlitz +3

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and ver…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.