Publications (16)
The Eval4NLP 2023 Shared Task on Prompting Large Language Models as Explainable Metrics
Christoph Leiter, Juri Opitz, Daniel Deutsch +3
With an increasing number of parameters and pre-training data, generative large language models (LLMs) have shown remarkable capabilities to solve tasks with minimal or no task-rel…
Human-in-the-Loop Schema Induction
Tianyi Zhang, Isaac Tham, Zhaoyi Hou +12
Schema induction builds a graph representation explaining how events unfold in a scenario. Existing approaches have been based on information retrieval (IR) and information extract…
A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods
Daniel Deutsch, Rotem Dror, Dan Roth
The quality of a summarization evaluation metric is quantified by calculating the correlation between its scores and human annotations across a large number of summaries. Currently…
State of What Art? A Call for Multi-Prompt LLM Evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin +3
Recent advances in large language models (LLMs) have led to the development of various evaluation benchmarks. These benchmarks typically rely on a single instruction template for e…
DMLR: Data-centric Machine Learning Research -- Past, Present and Future
Luis Oala, Manil Maskey, Lilith Bat-Leah +35
Drawing from discussions at the inaugural DMLR workshop at ICML 2023 and meetings prior, in this report we outline the relevance of community engagement and infrastructure developm…
Zero-Shot On-the-Fly Event Schema Induction
Rotem Dror, Haoyu Wang, Dan Roth
What are the events involved in a pandemic outbreak? What steps should be taken when planning a wedding? The answers to these questions can be found by collecting many documents on…
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
Nitay Calderon, Roi Reichart, Rotem Dror
The "LLM-as-an-annotator" and "LLM-as-a-judge" paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM a…
On the Limitations of Reference-Free Evaluations of Generated Text
Daniel Deutsch, Rotem Dror, Dan Roth
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can…
Diffusion Denoiser-Aided Gyrocompassing
Gershy Ben-Arie, Daniel Engelsman, Rotem Dror +1
The paper introduces a diffusion‑based denoising front‑end combined with a deep learning heading estimator to improve gyrocompassing accuracy for low‑cost gyroscopes, achieving not…
Appendix - Recommended Statistical Significance Tests for NLP Tasks
Rotem Dror, Roi Reichart
Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like t…
The Structured Weighted Violations Perceptron Algorithm
Rotem Dror, Roi Reichart
We present the Structured Weighted Violations Perceptron (SWVP) algorithm, a new structured prediction algorithm that generalizes the Collins Structured Perceptron (CSP). Unlike CS…
Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics
Daniel Deutsch, Rotem Dror, Dan Roth
How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which th…
Diffusion-Driven Inertial Generated Data for Smartphone Location Classification
Noa Cohen, Rotem Dror, Itzik Klein
Despite the crucial role of inertial measurements in motion tracking and navigation systems, the time-consuming and resource-intensive nature of collecting extensive inertial data…
Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
Rotem Dror, Gili Baumer, Marina Bogomolov +1
With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in orde…
AI Planning Framework for LLM-Based Web Agents
Orit Shahnovsky, Rotem Dror
Developing autonomous agents for web-based tasks is a core challenge in AI. While Large Language Model (LLM) agents can interpret complex user requests, they often operate as black…
The Structured Weighted Violations MIRA
Dor Ringel, Rotem Dror, Roi Reichart
We present the Structured Weighted Violation MIRA (SWVM), a new structured prediction algorithm that is based on an hybridization between MIRA (Crammer and Singer, 2003) and the st…