◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yotam Perlitz

17 papers hereh-index 9343 citations27 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author12
  • last author2

Across the 17 of 17 papers where every author was matched, so the position is known.

fields
  • cs.CL12
  • cs.AI4
  • cs.LG1
same name
  • Yotam Perlitz — 1 paper
  • Yotam Perlitz — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedGeneral Agent Evaluation

2 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.AIShow all

4 papers · 1 filter

cs.AI2026

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Tomer Keren, Nitay Calderon, Asaf Yehudai +3

As agent capabilities advance, existing benchmarks, such as I¨„2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and lab…

cs.AI2026★ 2 cited

General Agent Evaluation

Elron Bandel, Asaf Yehudai, Lilach Eden +12

General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes…

cs.AI2026

ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities

Christopher Zanoli, Andrea Giovannini, Tengjun Jin +2

Constructing Extract-Load-Transform (ELT) pipelines is a labor-intensive data engineering task and a high-impact target for AI automation. On ELT-Bench, the first benchmark for end…

cs.AI2025

How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability

Ora Nova Fandina, Leshem Choshen, Eitan Farchi +3

Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs,…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.