papers

Publications (73)

cs.CL2021

Data Efficient Masked Language Modeling for Vision and Language

Yonatan Bitton, Gabriel Stanovsky, Michael Elhadad +1

Masked language modeling (MLM) is one of the key sub-tasks in vision-language pretraining. In the cross-modal setting, tokens in the sentence are masked at random, and the model pr…

cs.CV2022

VASR: Visual Analogies of Situation Recognition

Yonatan Bitton, Ron Yosef, Eli Strugo +3

A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Anal…

cs.CL2017

Crowdsourcing Question-Answer Meaning Representations

Julian Michael, Gabriel Stanovsky, Luheng He +2

We introduce Question-Answer Meaning Representations (QAMRs), which represent the predicate-argument structure of a sentence as a set of question-answer pairs. We also develop a cr…

cs.CL2021

Process-Level Representation of Scientific Protocols with Interactive Annotation

Ronen Tamari, Fan Bai, Alan Ritter +1

We develop Process Execution Graphs (PEG), a document-level representation of real-world wet lab biochemistry protocols, addressing challenges such as cross-sentence relations, lon…

cs.CL2024

Schema-Driven Information Extraction from Heterogeneous Tables

Fan Bai, Junmo Kang, Gabriel Stanovsky +3

In this paper, we explore the question of whether large language models can support cost-efficient information extraction from tables. We introduce schema-driven information extrac…

cs.CL2023

You Can Have Your Data and Balance It Too: Towards Balanced and Efficient Multilingual Models

Tomasz Limisiewicz, Dan Malkin, Gabriel Stanovsky

Multilingual models have been widely used for cross-lingual transfer to low-resource languages. However, the performance on these languages is hindered by their underrepresentation…

cs.CL2025

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

Itay Itzhak, Yonatan Belinkov, Gabriel Stanovsky

Large language models (LLMs) exhibit cognitive biases -- systematic tendencies of irrational decision-making, similar to those seen in humans. Prior work has found that these biase…

cs.CL2023

Evaluating and Improving the Coreference Capabilities of Machine Translation Models

Asaf Yehudai, Arie Cattan, Omri Abend +1

Machine translation (MT) requires a wide range of linguistic capabilities, which current end-to-end models are expected to learn implicitly by observing aligned sentences in biling…

cs.CL2026

PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation

Eliya Habba, Noam Dahan, Gili Lior +1

Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for…

cs.CL2026

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Itay Itzhak, Eliya Habba, Gabriel Stanovsky +1

Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based ev…

cs.CL2022

WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language Models

Yonatan Bitton, Nitzan Bitton Guetta, Ron Yosef +4

While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we…

cs.CL2023

The Perfect Victim: Computational Analysis of Judicial Attitudes towards Victims of Sexual Violence

Eliya Habba, Renana Keydar, Dan Bareket +1

We develop computational models to analyze court statements in order to assess judicial attitudes toward victims of sexual violence in the Israeli court system. The study examines…

cs.CL2025

HACK: Hallucinations Along Certainty and Knowledge Axes

Adi Simhi, Jonathan Herzig, Itay Itzhak +7

Hallucinations in LLMs present a critical barrier to their reliable usage. Existing research usually categorizes hallucination by their external properties rather than by the LLMs'…

cs.CL2020

Active Learning for Coreference Resolution using Discrete Annotation

Belinda Z. Li, Gabriel Stanovsky, Luke Zettlemoyer

We improve upon pairwise annotation for active learning in coreference resolution, by asking annotators to identify mention antecedents if a presented mention pair is deemed not co…

cs.CL2025

Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages

Noam Dahan, Omer Kidron, Gabriel Stanovsky

High quality summarization data remains scarce in under-represented languages. However, historical newspapers, made available through recent digitization efforts, offer an abundant…

cs.CL2020

MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics

Anthony Chen, Gabriel Stanovsky, Sameer Singh +1

Posing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers. However, pro…

cs.CL2025

In-Context Learning on a Budget: A Case Study in Token Classification

Uri Berger, Tal Baumel, Gabriel Stanovsky

Few shot in-context learning (ICL) typically assumes access to large annotated training sets. However, in many real world scenarios, such as domain adaptation, there is only a limi…

cs.CL2022

A Computational Acquisition Model for Multimodal Word Categorization

Uri Berger, Gabriel Stanovsky, Omri Abend +1

Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition, which is believed to rely heavily on c…

cs.CL2021

Automated Extraction of Sentencing Decisions from Court Cases in the Hebrew Language

Mohr Wenger, Tom Kalir, Noga Berger +3

We present the task of Automated Punishment Extraction (APE) in sentencing decisions from criminal court cases in Hebrew. Addressing APE will enable the identification of sentencin…

cs.CL2025

The State and Fate of Summarization Datasets: A Survey

Noam Dahan, Gabriel Stanovsky

Automatic summarization has consistently attracted attention due to its versatility and wide application in various downstream tasks. Despite its popularity, we find that annotatio…

cs.CL2021

Cross-document Coreference Resolution over Predicted Mentions

Arie Cattan, Alon Eirew, Gabriel Stanovsky +2

Coreference resolution has been mostly investigated within a single document scope, showing impressive progress in recent years based on end-to-end models. However, the more challe…

cs.CL2020

Controlled Crowdsourcing for High-Quality QA-SRL Annotation

Paul Roit, Ayal Klein, Daniela Stepanov +5

Question-answer driven Semantic Role Labeling (QA-SRL) was proposed as an attractive open and natural flavour of SRL, potentially attainable from laymen. Recently, a large-scale cr…

cs.CL2024

Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition

Ariel Goldstein, Gabriel Stanovsky

Recent advances in LLMs have sparked a debate on whether they understand text. In this position paper, we argue that opponents in this debate hold different definitions for underst…

cs.CL2020

Streamlining Cross-Document Coreference Resolution: Evaluation and Modeling

Arie Cattan, Alon Eirew, Gabriel Stanovsky +2

Recent evaluation protocols for Cross-document (CD) coreference resolution have often been inconsistent or lenient, leading to incomparable results across works and overestimation…

cs.CL2021

Collecting a Large-Scale Gender Bias Dataset for Coreference Resolution and Machine Translation

Shahar Levy, Koren Lazar, Gabriel Stanovsky

Recent works have found evidence of gender bias in models of machine translation and coreference resolution using mostly synthetic diagnostic datasets. While these quantify bias in…

cs.CL2026

Anticipatory Evaluation of Language Models

Jungsoo Park, Ethan Mendes, Gabriel Stanovsky +1

Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whethe…

cs.CL2023

Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution

Gili Lior, Gabriel Stanovsky

Spurious correlations were found to be an important factor explaining model performance in various NLP tasks (e.g., gender or racial artifacts), often considered to be ''shortcuts'…

cs.CL2025

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments

Gili Lior, Eliya Habba, Shahar Levy +2

LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations…

cs.CL2019

Evaluating Gender Bias in Machine Translation

Gabriel Stanovsky, Noah A. Smith, Luke Zettlemoyer

We present the first challenge set and evaluation protocol for the analysis of gender bias in machine translation (MT). Our approach uses two recent coreference resolution datasets…

cs.CL2026

When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking

Orian Dabod, Amir DN Cohen, Gabriel Stanovsky

Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying that the expensive reranking step can in f…

cs.LG2025

Beyond Benchmarks: On The False Promise of AI Regulation

Gabriel Stanovsky, Renana Keydar, Gadi Perl +1

The performance of AI models on safety benchmarks does not indicate their real-world performance after deployment. This opaqueness of AI models impedes existing regulatory framewor…

cs.CL2025

Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer

Adi Simhi, Itay Itzhak, Fazl Barez +2

Prior work on large language model (LLM) hallucinations has associated them with model uncertainty or inaccurate knowledge. In this work, we define and investigate a distinct type…

cs.CL2026

Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration

Eliya Habba, Itay Itzhak, Asaf Yehudai +5

The rapid release of both language models and benchmarks makes it increasingly costly to evaluate every model on every dataset. In practice, models are often evaluated on different…

cs.CL2022

GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation

Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg +5

While often assumed a gold standard, effective human evaluation of text generation remains an important, open area for research. We revisit this problem with a focus on producing c…

cs.CL2026

DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation

Eliya Habba, Ofir Arviv, Itay Itzhak +5

Recent work found that LLMs are sensitive to a wide range of arbitrary prompt dimensions, including the type of delimiters, answer enumerators, instruction wording, and more. This…

cs.CL2024

Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation

Bar Iluz, Yanai Elazar, Asaf Yehudai +1

Most works on gender bias focus on intrinsic bias -- removing traces of information about a protected group from the model's internal representation. However, these works are often…

cs.CL2025

Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis

Uri Berger, Gabriel Stanovsky, Omri Abend +1

The task of image captioning has recently been gaining popularity, and with it the complex task of evaluating the quality of image captioning models. In this work, we present the f…

cs.CL2024

State of What Art? A Call for Multi-Prompt LLM Evaluation

Moran Mizrahi, Guy Kaplan, Dan Malkin +3

Recent advances in large language models (LLMs) have led to the development of various evaluation benchmarks. These benchmarks typically rely on a single instruction template for e…

cs.CL2019

Yall should read this! Identifying Plurality in Second-Person Personal Pronouns in English Texts

Gabriel Stanovsky, Ronen Tamari

Distinguishing between singular and plural "you" in English is a challenging task which has potential for downstream applications, such as machine translation or coreference resolu…

cs.CL2024

SAUCE: Synchronous and Asynchronous User-Customizable Environment for Multi-Agent LLM Interaction

Shlomo Neuberger, Niv Eckhaus, Uri Berger +3

Many human interactions, such as political debates, are carried out in group settings, where there are arbitrarily many participants, each with different views and agendas. To expl…

cs.CL2022

On the Limitations of Dataset Balancing: The Lost Battle Against Spurious Correlations

Roy Schwartz, Gabriel Stanovsky

Recent work has shown that deep learning models in NLP are highly sensitive to low-level correlations between simple features and specific output labels, leading to overfitting and…

cs.CL2024

Looking Beyond The Top-1: Transformers Determine Top Tokens In Order

Daria Lioubashevski, Tomer Schlank, Gabriel Stanovsky +1

Understanding the inner workings of Transformers is crucial for achieving more accurate and efficient predictions. In this work, we analyze the computation performed by Transformer…

cs.CL2021

Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQA

Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz +1

Recent works have shown that supervised models often exploit data artifacts to achieve good test scores while their performance severely degrades on samples outside their training…

cs.CL2019

On the Limits of Learning to Actively Learn Semantic Representations

Omri Koshorek, Gabriel Stanovsky, Yichu Zhou +2

One of the goals of natural language understanding is to develop models that map sentences into meaning representations. However, training such models requires expensive annotation…

cs.CL2022

"Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun Compounds

Keshav Kolluru, Gabriel Stanovsky, Mausam

Proper noun compounds, e.g., "Covid vaccine", convey information in a succinct manner (a "Covid vaccine" is a "vaccine that immunizes against the Covid disease"). These are commonl…

cs.AI2020

Ecological Semantics: Programming Environments for Situated Language Understanding

Ronen Tamari, Gabriel Stanovsky, Dafna Shahaf +1

Large-scale natural language understanding (NLU) systems have made impressive progress: they can be applied flexibly across a variety of tasks, and employ minimal structural assump…

cs.MA2025

Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games

Niv Eckhaus, Uri Berger, Gabriel Stanovsky

LLMs are used predominantly in synchronous communication, where a human user and a model communicate in alternating turns. In contrast, many real-world settings are asynchronous. F…

cs.CL2025

Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs

Jungsoo Park, Junmo Kang, Gabriel Stanovsky +1

The surge of LLM studies makes synthesizing their findings challenging. Analysis of experimental results from literature can uncover important trends across studies, but the time-c…

cs.CL2023

Are Layout-Infused Language Models Robust to Layout Distribution Shifts? A Case Study with Scientific Documents

Catherine Chen, Zejiang Shen, Dan Klein +3

Recent work has shown that infusing layout features into language models (LMs) improves processing of visually-rich documents such as scientific papers. Layout-infused LMs are ofte…

cs.CL2026

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…

cs.CL2026

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

Shahar Levy, Eliya Habba, Reshef Mintz +3

Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually de…

cs.CL2023

Exploring the Impact of Training Data Distribution and Subword Tokenization on Gender Bias in Machine Translation

Bar Iluz, Tomasz Limisiewicz, Gabriel Stanovsky +1

We study the effect of tokenization on gender bias in machine translation, an aspect that has been largely overlooked in previous works. Specifically, we focus on the interactions…

cs.CL2025

More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG

Shahar Levy, Nir Mazor, Lihi Shalmon +2

Retrieval-Augmented Generation (RAG) enhances the accuracy of Large Language Model (LLM) responses by leveraging relevant external documents during generation. Although previous st…

cs.CL2026

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

Reef Menaged, Gili Lior, Shauli Ravfogel +2

We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…

cs.CL2024

SEAM: A Stochastic Benchmark for Multi-Document Tasks

Gili Lior, Avi Caciularu, Arie Cattan +3

Various tasks, such as summarization, multi-hop question answering, or coreference resolution, are naturally phrased over collections of real-world documents. Such tasks present a…

cs.CV2025

Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time

Uri Berger, Omri Abend, Lea Frermann +1

Incorporating automatically predicted human feedback into the process of training generative models has attracted substantial recent interest, while feedback at inference time has…

cs.DL2021

Gender trends in computer science authorship

Lucy Lu Wang, Gabriel Stanovsky, Luca Weihs +1

A large-scale, up-to-date analysis of Computer Science literature (11.8M papers through 2019) reveals that, if trends from the last 50 years continue, parity between the number of…

cs.CL2024

Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction

Gili Lior, Yoav Goldberg, Gabriel Stanovsky

Document collections of various domains, e.g., legal, medical, or financial, often share some underlying collection-wide structure, which captures information that can aid both hum…

cs.CL2025

Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination

Moran Mizrahi, Chen Shani, Gabriel Stanovsky +2

Large Language Models (LLMs) excel at many tasks, yet they struggle to produce truly creative, diverse ideas. In this paper, we introduce a novel approach that enhances LLM creativ…

cs.CL2026

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

Adi Gabay, Gabriel Stanovsky, Liat Peterfreund

Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior work evaluating LLMs on epistemic…

cs.CL2020

The Right Tool for the Job: Matching Model and Instance Complexities

Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta +2

As NLP models become larger, executing a trained model requires significant computational resources incurring monetary and environmental costs. To better respect a given inference…

cs.CV2023

Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images

Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel +4

Weird, unusual, and uncanny images pique the curiosity of observers because they challenge commonsense. For example, an image released during the 2022 world cup depicts the famous…

cs.CL2026

Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts

Gili Lior, Liron Nacchace, Gabriel Stanovsky

Humans are influenced by how information is presented, a phenomenon known as the framing effect. Prior work suggests that LLMs may also be susceptible to framing, but it has relied…

cs.CL2024

A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns

Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky +2

Cross-domain alignment refers to the task of mapping a concept from one domain to another. For example, ``If a \textit{doctor} were a \textit{color}, what color would it be?''. Thi…

cs.CL2024

K-QA: A Real-World Medical Q&A Benchmark

Itay Manes, Naama Ronn, David Cohen +3

Ensuring the accuracy of responses provided by large language models (LLMs) is crucial, particularly in clinical settings where incorrect information may directly impact patient he…

cs.CL2022

A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank

Dan Malkin, Tomasz Limisiewicz, Gabriel Stanovsky

We show that the choice of pretraining languages affects downstream cross-lingual transfer for BERT-based models. We inspect zero-shot performance in balanced data conditions to mi…

cs.CL2021

Realistic Evaluation Principles for Cross-document Coreference Resolution

Arie Cattan, Alon Eirew, Gabriel Stanovsky +2

We point out that common evaluation practices for cross-document coreference resolution have been unrealistically permissive in their assumed settings, yielding inflated results. W…

cs.CL2023

A Large-Scale Multilingual Study of Visual Constraints on Linguistic Selection of Descriptions

Uri Berger, Lea Frermann, Gabriel Stanovsky +1

We present a large, multilingual study into how vision constrains linguistic choice, covering four languages and five linguistic properties, such as verb transitivity or use of num…

cs.AI2024

Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias

Itay Itzhak, Gabriel Stanovsky, Nir Rosenfeld +1

Recent studies show that instruction tuning (IT) and reinforcement learning from human feedback (RLHF) improve the abilities of large language models (LMs) dramatically. While thes…

cs.CL2019

DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Dheeru Dua, Yizhong Wang, Pradeep Dasigi +3

Reading comprehension has recently seen rapid progress, with systems matching humans on the most popular datasets for the task. However, a large body of work has highlighted the br…

cs.CL2021

Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling Approach

Koren Lazar, Benny Saret, Asaf Yehudai +3

We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE - 100 CE). Due to the…

cs.CL2020

Gender Coreference and Bias Evaluation at WMT 2020

Tom Kocmi, Tomasz Limisiewicz, Gabriel Stanovsky

Gender bias in machine translation can manifest when choosing gender inflections based on spurious gender correlations. For example, always translating doctors as men and nurses as…

cs.CL2016

Getting More Out Of Syntax with PropS

Gabriel Stanovsky, Jessica Ficler, Ido Dagan +1

Semantic NLP applications often rely on dependency trees to recognize major elements of the proposition structure of sentences. Yet, while much semantic structure is indeed express…