papers

Publications (38)

cs.CL2021

Blindness to Modality Helps Entailment Graph Mining

Liane Guillou, Sander Bijl de Vroe, Mark Johnson +1

Understanding linguistic modality is widely seen as important for downstream tasks such as Question Answering and Knowledge Graph Population. Entailment Graph learning might also b…

cs.CL2019

Data Augmentation via Dependency Tree Morphing for Low-Resource Languages

Gözde Gül Şahin, Mark Steedman

Neural NLP systems achieve high scores in the presence of sizable training dataset. Lack of such datasets leads to poor system performances in the case low-resource languages. We p…

cs.CL2020

The role of context in neural pitch accent detection in English

Elizabeth Nielsen, Mark Steedman, Sharon Goldwater

Prosody is a rich information source in natural language, serving as a marker for phenomena such as contrast. In order to make this information available to downstream tasks, we ne…

cs.CL2022

Zero-shot Cross-Linguistic Learning of Event Semantics

Malihe Alikhani, Thomas Kober, Bashar Alhafni +6

Typologically diverse languages offer systems of lexical and grammatical aspect that allow speakers to focus on facets of event structure in ways that comport with the specific com…

cs.CL2025

S2LPP: Small-to-Large Prompt Prediction across LLMs

Liang Cheng, Tianyi LI, Zhaowei Wang +1

The performance of pre-trained Large Language Models (LLMs) is often sensitive to nuances in prompt templates, requiring careful prompt engineering, adding costs in terms of comput…

cs.CL2023

Smoothing Entailment Graphs with Language Models

Nick McKenna, Tianyi Li, Mark Johnson +1

The diversity and Zipfian frequency distribution of natural language predicates in corpora leads to sparsity in Entailment Graphs (EGs) built by Open Relation Extraction (ORE). EGs…

cs.LG2012

Learning STRIPS Operators from Noisy and Incomplete Observations

Kira Mourao, Luke S. Zettlemoyer, Ronald P. A. Petrick +1

Agents learning to act autonomously in real-world domains must acquire a model of the dynamics of the domain in which they operate. Learning domain dynamics can be challenging, esp…

cs.CL2023

Modeling structure-building in the brain with CCG parsing and large language models

Miloš Stanojević, Jonathan R. Brennan, Donald Dunagan +2

To model behavioral and neural correlates of language comprehension in naturalistic environments researchers have turned to broad-coverage tools from natural-language processing an…

cs.CL2017

Evaluating Induced CCG Parsers on Grounded Semantic Parsing

Yonatan Bisk, Siva Reddy, John Blitzer +2

We compare the effectiveness of four different syntactic CCG parsers for a semantic slot-filling task to explore how much syntactic supervision is required for downstream semantic…

cs.CL2022

Jointly Modeling Hierarchical and Horizontal Features for Relational Triple Extraction

Zhepei Wei, Yantao Jia, Yuan Tian +4

Recent works on relational triple extraction have shown the superiority of jointly extracting entities and relations over the pipelined extraction manner. However, most existing jo…

cs.CL2024

A Usage-centric Take on Intent Understanding in E-Commerce

Wendi Zhou, Tianyi Li, Pavlos Vougiouklis +2

Identifying and understanding user intents is a pivotal task for E-Commerce. Despite its essential role in product recommendation and business user profiling analysis, intent under…

cmp-lg1994

Specifying Intonation from Context for Speech Synthesis

Scott Prevost, Mark Steedman

This paper presents a theory and a computational implementation for generating prosodically appropriate synthetic speech in response to database queries. Proper distinctions of con…

cs.CV2025

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Zhaowei Wang, Wenhao Yu, Xiyu Ren +9

The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of ima…

cs.CL2021

Incorporating Temporal Information in Entailment Graph Mining

Liane Guillou, Sander Bijl de Vroe, Mohammad Javad Hosseini +2

We present a novel method for injecting temporality into entailment graphs to address the problem of spurious entailments, which may arise from similar but temporally distinct even…

cs.CL2024

Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets

Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5

Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…

cs.CL2022

Cross-lingual Inference with A Chinese Entailment Graph

Tianyi Li, Sabine Weber, Mohammad Javad Hosseini +2

Predicate entailment detection is a crucial task for question-answering from text, where previous work has explored unsupervised learning of entailment graphs from typed open relat…

cs.CL2023

Sentence-Incremental Neural Coreference Resolution

Matt Grenander, Shay B. Cohen, Mark Steedman

We propose a sentence-incremental neural coreference resolution system which incrementally builds clusters after marking mention boundaries in a shift-reduce method. The system is…

cs.CL2025

Neutralizing Bias in LLM Reasoning using Entailment Graphs

Liang Cheng, Tianyi Li, Zhaowei Wang +2

LLMs are often claimed to be capable of Natural Language Inference (NLI), which is widely regarded as a cornerstone of more complex forms of reasoning. However, recent works show t…

cs.CL2024

Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction

Kaiqiao Han, Tianqing Fang, Zhaowei Wang +2

While Large Language Models (LLMs) have showcased remarkable proficiency in reasoning, there is still a concern about hallucinations and unreliable reasoning issues due to semantic…

cs.CL2021

Cross-lingual Intermediate Fine-tuning improves Dialogue State Tracking

Nikita Moghe, Mark Steedman, Alexandra Birch

Recent progress in task-oriented neural dialogue systems is largely focused on a handful of languages, as annotation of training data is tedious and expensive. Machine translation…

cs.CL2025

LLMs are Frequency Pattern Learners in Natural Language Inference

Liang Cheng, Zhaowei Wang, Mark Steedman

While fine-tuning LLMs on NLI corpora improves their inferential performance, the underlying mechanisms driving this improvement remain largely opaque. In this work, we conduct a s…

cs.CL2024

Explicit Inductive Inference using Large Language Models

Tianyang Liu, Tianyi Li, Liang Cheng +1

Large Language Models (LLMs) are reported to hold undesirable attestation bias on inference tasks: when asked to predict if a premise P entails a hypothesis H, instead of consideri…

cs.CL2022

Language Models Are Poor Learners of Directional Inference

Tianyi Li, Mohammad Javad Hosseini, Sabine Weber +1

We examine LMs' competence of directional predicate entailments by supervised fine-tuning with prompts. Our analysis shows that contrary to their apparent success on standard NLI,…

cs.CL2024

Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech

Ida Szubert, Omri Abend, Nathan Schneider +4

This paper proposes a methodology for constructing such corpora of child directed speech (CDS) paired with sentential logical forms, and uses this method to create two such corpora…

cs.CL2023

Extrinsic Evaluation of Machine Translation Metrics

Nikita Moghe, Tom Sherborne, Mark Steedman +1

Automatic machine translation (MT) metrics are widely used to distinguish the translation qualities of machine translation systems across relatively large test sets (system-level e…

cs.CL2023

Multi-Document Summarization with Centroid-Based Pretraining

Ratish Puduppully, Parag Jain, Nancy F. Chen +1

In Multi-Document Summarization (MDS), the input can be modeled as a set of documents, and the output is its summary. In this paper, we focus on pretraining objectives for MDS. Spe…

cs.CL2023

Prosodic features improve sentence segmentation and parsing

Elizabeth Nielsen, Sharon Goldwater, Mark Steedman

Parsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries. We know from previous work that prosody helps in parsing sing…

cs.CL2021

Prosodic segmentation for parsing spoken dialogue

Elizabeth Nielsen, Mark Steedman, Sharon Goldwater

Parsing spoken dialogue poses unique difficulties, including disfluencies and unmarked boundaries between sentence-like units. Previous work has shown that prosody can help with pa…

cs.CL2024

A Language-agnostic Model of Child Language Acquisition

Louis Mahon, Omri Abend, Uri Berger +3

This work reimplements a recent semantic bootstrapping child-language acquisition model, which was originally designed for English, and trains it to learn a new language: Hebrew. T…

cs.CL2025

Modelling Child Learning and Parsing of Long-range Syntactic Dependencies

Louis Mahon, Mark Johnson, Mark Steedman

This work develops a probabilistic child language acquisition model to learn a range of linguistic phenonmena, most notably long-range syntactic dependencies of the sort found in o…

cs.CL2025

Efficient Seq2seq Coreference Resolution Using Entity Representations

Matt Grenander, Shay B. Cohen, Mark Steedman

Seq2seq coreference models have introduced a new paradigm for coreference resolution by learning to generate text corresponding to coreference labels, without requiring task-specif…

cs.CL2017

Universal Semantic Parsing

Siva Reddy, Oscar Täckström, Slav Petrov +2

Universal Dependencies (UD) offer a uniform cross-lingual syntactic representation, with the aim of advancing multilingual applications. Recent work shows that semantic parsing can…

cs.CL2021

Multivalent Entailment Graphs for Question Answering

Nick McKenna, Liane Guillou, Mohammad Javad Hosseini +3

Drawing inferences between open-domain natural language predicates is a necessity for true language understanding. There has been much progress in unsupervised learning of entailme…

cs.CL2019

Temporal and Aspectual Entailment

Thomas Kober, Sander Bijl de Vroe, Mark Steedman

Inferences regarding "Jane's arrival in London" from predications such as "Jane is going to London" or "Jane has gone to London" depend on tense and aspect of the predications. Ten…

cs.CL2021

Modality and Negation in Event Extraction

Sander Bijl de Vroe, Liane Guillou, Miloš Stanojević +2

Language provides speakers with a rich system of modality for expressing thoughts about events, without being committed to their actual occurrence. Modality is commonly used in the…

cs.CL2020

Aspectuality Across Genre: A Distributional Semantics Approach

Thomas Kober, Malihe Alikhani, Matthew Stone +1

The interpretation of the lexical aspect of verbs in English plays a crucial role for recognizing textual entailment and learning discourse-level inferences. We show that two eleme…

cs.CL2023

Sources of Hallucination by Large Language Models on Inference Tasks

Nick McKenna, Tianyi Li, Liang Cheng +3

Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI), necessary for applied tasks like question answering and summarization. We present a seri…

cs.CL2018

Character-Level Models versus Morphology in Semantic Role Labeling

Gözde Gül Şahin, Mark Steedman

Character-level models have become a popular approach specially for their accessibility and ability to handle unseen data. However, little is known on their ability to reveal the u…