papers

Publications (16)

cs.CL2023

The Eval4NLP 2023 Shared Task on Prompting Large Language Models as Explainable Metrics

Christoph Leiter, Juri Opitz, Daniel Deutsch +3

With an increasing number of parameters and pre-training data, generative large language models (LLMs) have shown remarkable capabilities to solve tasks with minimal or no task-rel…

cs.HC2023

Human-in-the-Loop Schema Induction

Tianyi Zhang, Isaac Tham, Zhaoyi Hou +12

Schema induction builds a graph representation explaining how events unfold in a scenario. Existing approaches have been based on information retrieval (IR) and information extract…

cs.CL2021

A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods

Daniel Deutsch, Rotem Dror, Dan Roth

The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently…

cs.CL2024

State of What Art? A Call for Multi-Prompt LLM Evaluation

Moran Mizrahi, Guy Kaplan, Dan Malkin +3

Recent advances in large language models (LLMs) have led to the development of various evaluation benchmarks. These benchmarks typically rely on a single instruction template for e…

cs.LG2024

DMLR: Data-centric Machine Learning Research -- Past, Present and Future

Luis Oala, Manil Maskey, Lilith Bat-Leah +35

Drawing from discussions at the inaugural DMLR workshop at ICML 2023 and meetings prior, in this report we outline the relevance of community engagement and infrastructure developm…

cs.CL2023

Zero-Shot On-the-Fly Event Schema Induction

Rotem Dror, Haoyu Wang, Dan Roth

What are the events involved in a pandemic outbreak? What steps should be taken when planning a wedding? The answers to these questions can be found by collecting many documents on…

cs.CL2025

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

Nitay Calderon, Roi Reichart, Rotem Dror

The "LLM-as-an-annotator" and "LLM-as-a-judge" paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM a…

cs.CL2022

On the Limitations of Reference-Free Evaluations of Generated Text

Daniel Deutsch, Rotem Dror, Dan Roth

There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can…

cs.RO2026

Diffusion Denoiser-Aided Gyrocompassing

Gershy Ben-Arie, Daniel Engelsman, Rotem Dror +1

The paper introduces a diffusion‑based denoising front‑end combined with a deep learning heading estimator to improve gyrocompassing accuracy for low‑cost gyroscopes, achieving not…

#gyrocompassing#diffusion denoising#inertial navigation#low-cost sensors
cs.CL2018

Appendix - Recommended Statistical Significance Tests for NLP Tasks

Rotem Dror, Roi Reichart

Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like t…

cs.LG2016

The Structured Weighted Violations Perceptron Algorithm

Rotem Dror, Roi Reichart

We present the Structured Weighted Violations Perceptron (SWVP) algorithm, a new structured prediction algorithm that generalizes the Collins Structured Perceptron (CSP). Unlike CS…

cs.CL2022

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Daniel Deutsch, Rotem Dror, Dan Roth

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which th…

cs.LG2025

Diffusion-Driven Inertial Generated Data for Smartphone Location Classification

Noa Cohen, Rotem Dror, Itzik Klein

Despite the crucial role of inertial measurements in motion tracking and navigation systems, the time-consuming and resource-intensive nature of collecting extensive inertial data…

cs.CL2017

Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets

Rotem Dror, Gili Baumer, Marina Bogomolov +1

With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in orde…

cs.AI2026

AI Planning Framework for LLM-Based Web Agents

Orit Shahnovsky, Rotem Dror

Developing autonomous agents for web-based tasks is a core challenge in AI. While Large Language Model (LLM) agents can interpret complex user requests, they often operate as black…

cs.CL2020

The Structured Weighted Violations MIRA

Dor Ringel, Rotem Dror, Roi Reichart

We present the Structured Weighted Violation MIRA (SWVM), a new structured prediction algorithm that is based on an hybridization between MIRA (Crammer and Singer, 2003) and the st…