papers

Publications (151)

cs.CL2020

What Can We Do to Improve Peer Review in NLP?

Anna Rogers, Isabelle Augenstein

Peer review is our best tool for judging the quality of conference submissions, but it is becoming increasingly spurious. We argue that a part of the problem is that the reviewers…

cs.CL2018

Multi-task Learning of Pairwise Sequence Classification Tasks Over Disparate Label Spaces

Isabelle Augenstein, Sebastian Ruder, Anders Søgaard

We combine multi-task learning and semi-supervised learning by inducing a joint embedding space between disparate label spaces and learning transfer functions between label embeddi…

cs.AI2021

Disembodied Machine Learning: On the Illusion of Objectivity in NLP

Zeerak Waseem, Smarika Lulz, Joachim Bingel +1

Machine Learning seeks to identify and encode bodies of knowledge within provided datasets. However, data encodes subjective content, which determines the possible outcomes of the…

cs.CL2018

Copenhagen at CoNLL--SIGMORPHON 2018: Multilingual Inflection in Context with Explicit Morphosyntactic Decoding

Yova Kementchedjhieva, Johannes Bjerva, Isabelle Augenstein

This paper documents the Team Copenhagen system which placed first in the CoNLL--SIGMORPHON 2018 shared task on universal morphological reinflection, Task 2 with an overall accurac…

cs.CL2018

A simple but tough-to-beat baseline for the Fake News Challenge stance detection task

Benjamin Riedel, Isabelle Augenstein, Georgios P. Spithourakis +1

Identifying public misinformation is a complicated and challenging task. An important part of checking the veracity of a specific claim is to evaluate the stance different news sou…

cs.CL2022

Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection

Indira Sen, Mattia Samory, Claudia Wagner +1

Counterfactually Augmented Data (CAD) aims to improve out-of-domain generalizability, an indicator of model robustness. The improvement is credited with promoting core features of…

cs.CL2017

Multi-Task Learning of Keyphrase Boundary Classification

Isabelle Augenstein, Anders Søgaard

Keyphrase boundary classification (KBC) is the task of detecting keyphrases in scientific articles and labelling them with respect to predefined types. Although important in practi…

cs.CL2026

BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation

Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3

Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…

cs.CL2021

Does Typological Blinding Impede Cross-Lingual Sharing?

Johannes Bjerva, Isabelle Augenstein

Bridging the performance gap between high- and low-resource languages has been the focus of much previous work. Typological features from databases such as the World Atlas of Langu…

cs.CL2026

Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns

Amalie Brogaard Pauli, Maria Barrett, Max Müller-Eberstein +2

Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade. Prior work has sh…

cs.CV2026

Interpretable Debiasing of Vision-Language Models for Social Fairness

Na Min An, Yoonna Jang, Yusuke Hirota +3

The rapid advancement of Vision-Language models (VLMs) has raised growing concerns that their black-box reasoning processes could lead to unintended forms of social bias. Current d…

cs.CL2021

CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding

Dustin Wright, Isabelle Augenstein

Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annota…

cs.CL2026

Output Vector Editing for Memorization Mitigation in Large Language Models

Ahmad Dawar Hakimi, Kaiwei Lei, Isabelle Augenstein +1

Large language models memorize and reproduce sequences from their training data, creating privacy, copyright, and security risks. Existing neuron-level mitigation methods equate ed…

cs.SI2025

Investigating Human Values in Online Communities

Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee +1

Studying human values is instrumental for cross-cultural research, enabling a better understanding of preferences and behaviour of society at large and communities therein. To stud…

cs.CL2016

Numerically Grounded Language Models for Semantic Error Correction

Georgios P. Spithourakis, Isabelle Augenstein, Sebastian Riedel

Semantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction. Current approaches generally fo…

cs.LG2020

TX-Ray: Quantifying and Explaining Model-Knowledge Transfer in (Un-)Supervised NLP

Nils Rethmeier, Vageesh Kumar Saxena, Isabelle Augenstein

While state-of-the-art NLP explainability (XAI) methods focus on explaining per-sample decisions in supervised end or probing tasks, this is insufficient to explain and quantify mo…

cs.CL2022

Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models

Karolina Stańczak, Edoardo Ponti, Lucas Torroba Hennigen +2

The success of multilingual pre-trained models is underpinned by their ability to learn representations shared by multiple languages even in absence of any explicit supervision. Ho…

cs.CL2018

From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings

Johannes Bjerva, Isabelle Augenstein

A core part of linguistic typology is the classification of languages according to linguistic properties, such as those detailed in the World Atlas of Language Structure (WALS). Do…

cs.CL2026

Epistemic Diversity and Knowledge Collapse in Large Language Models

Dustin Wright, Sarah Masud, Jared Moore +5

Large language models (LLMs) tend to generate homogenous texts, which may impact the diversity of knowledge generated across different outputs. Given their potential to replace exi…

cs.CL2024

People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection

Indira Sen, Dennis Assenmacher, Mattia Samory +3

NLP models are used in a variety of critical social computing tasks, such as detecting sexist, racist, or otherwise hateful content. Therefore, it is imperative that these models a…

cs.CL2025

A Reality Check on Context Utilisation for Retrieval-Augmented Generation

Lovisa Hagström, Sara Vera Marjanović, Haeun Yu +5

Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can…

cs.SI2025

Community Moderation and the New Epistemology of Fact Checking on Social Media

Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty +13

Social media platforms have traditionally relied on internal moderation teams and partnerships with independent fact-checking organizations to identify and flag misleading content.…

cs.CL2021

Inducing Language-Agnostic Multilingual Representations

Wei Zhao, Steffen Eger, Johannes Bjerva +1

Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world. However, they currently require large pretraining…

cs.CL2025

Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics

Amalie Brogaard Pauli, Isabelle Augenstein, Ira Assent

Large language models (LLMs) make it easy to rewrite a text in any style -- e.g. to make it more polite, persuasive, or more positive -- but evaluation thereof is not straightforwa…

cs.CL2021

Joint Emotion Label Space Modelling for Affect Lexica

Luna De Bruyne, Pepa Atanasova, Isabelle Augenstein

Emotion lexica are commonly used resources to combat data poverty in automatic emotion detection. However, vocabulary coverage issues, differences in construction method and discre…

cs.CL2020

SIGTYP 2020 Shared Task: Prediction of Typological Features

Johannes Bjerva, Elizabeth Salesky, Sabrina J. Mielke +6

Typological knowledge bases (KBs) such as WALS (Dryer and Haspelmath, 2013) contain information about linguistic properties of the world's languages. They have been shown to be use…

cs.CL2020

Zero-Shot Cross-Lingual Transfer with Meta Learning

Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva +1

Learning what to share between tasks has been a topic of great importance recently, as strategic sharing of knowledge has been shown to improve downstream task performance. This is…

cs.CL2022

A Neighbourhood Framework for Resource-Lean Content Flagging

Sheikh Muhammad Sarwar, Dimitrina Zlatkova, Momchil Hardalov +3

We propose a novel framework for cross-lingual content flagging with limited target-language data, which significantly outperforms prior work in terms of predictive performance. Th…

cs.CL2026

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein

Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understan…

cs.CL2021

Semi-Supervised Exaggeration Detection of Health Science Press Releases

Dustin Wright, Isabelle Augenstein

Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scienti…

cs.CL2021

Cross-Domain Label-Adaptive Stance Detection

Momchil Hardalov, Arnav Arora, Preslav Nakov +1

Stance detection concerns the classification of a writer's viewpoint towards a target. There are different task variants, e.g., stance of a tweet vs. a full article, or stance with…

cs.CL2020

Unsupervised Evaluation for Question Answering with Transformers

Lukas Muttenthaler, Isabelle Augenstein, Johannes Bjerva

It is challenging to automatically evaluate the answer of a QA model at inference time. Although many models provide confidence scores, and simple heuristics can go a long way towa…

cs.CL2023

Faithfulness Tests for Natural Language Explanations

Pepa Atanasova, Oana-Maria Camburu, Christina Lioma +3

Explanations of neural models aim to reveal a model's decision-making process for its predictions. However, recent work shows that current methods giving explanations such as salie…

cs.CL2025

Graph-Guided Textual Explanation Generation Framework

Shuzhou Yuan, Jingyi Sun, Ran Zhang +4

Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions. However, recent work has questioned…

cs.CL2025

Revealing Fine-Grained Values and Opinions in Large Language Models

Dustin Wright, Arnav Arora, Nadav Borenstein +3

Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…

cs.CL2024

The Causal Influence of Grammatical Gender on Distributional Semantics

Karolina Stańczak, Kevin Du, Adina Williams +2

How much meaning influences gender assignment across languages is an active area of research in linguistics and cognitive science. We can view current approaches as aiming to deter…

cs.CL2022

Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings

Malte Ostendorff, Nils Rethmeier, Isabelle Augenstein +2

Learning scientific document representations can be substantially improved through contrastive learning objectives, where the challenge lies in creating positive and negative train…

cs.CL2017

Discourse-Aware Rumour Stance Classification in Social Media Using Sequential Classifiers

Arkaitz Zubiaga, Elena Kochkina, Maria Liakata +5

Rumour stance classification, defined as classifying the stance of specific social media posts into one of supporting, denying, querying or commenting on an earlier post, is becomi…

cs.CL2019

Transductive Auxiliary Task Self-Training for Neural Multi-Task Models

Johannes Bjerva, Katharina Kann, Isabelle Augenstein

Multi-task learning and self-training are two common ways to improve a machine learning model's performance in settings with limited training data. Drawing heavily on ideas from th…

cs.HC2025

Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking

Greta Warren, Irina Shklovski, Isabelle Augenstein

The pervasiveness of large language models and generative AI in online media has amplified the need for effective automated fact-checking to assist fact-checkers in tackling the in…

cs.CL2022

Multi-Sense Language Modelling

Andrea Lekkas, Peter Schneider-Kamp, Isabelle Augenstein

The effectiveness of a language model is influenced by its token representations, which must encode contextual information and handle the same word form having a plurality of meani…

cs.CL2019

Uncovering Probabilistic Implications in Typological Knowledge Bases

Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell +1

The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have post…

cs.CL2019

Retrieval-based Goal-Oriented Dialogue Generation

Ana Valeria Gonzalez, Isabelle Augenstein, Anders Søgaard

Most research on dialogue has focused either on dialogue generation for openended chit chat or on state tracking for goal-directed dialogue. In this work, we explore a hybrid appro…

cs.CL2025

Quantifying Gender Biases Towards Politicians on Reddit

Sara Marjanovic, Karolina Stańczak, Isabelle Augenstein

Despite attempts to increase gender parity in politics, global efforts have struggled to ensure equal female representation. This is likely tied to implicit gender biases against w…

cs.CL2022

Modeling Information Change in Science Communication with Semantically Matched Paraphrases

Dustin Wright, Jiaxin Pei, David Jurgens +1

Whether the media faithfully communicate scientific information has long been a core issue to the science community. Automatically identifying paraphrased scientific findings could…

cs.CL2017

Generalisation in Named Entity Recognition: A Quantitative Analysis

Isabelle Augenstein, Leon Derczynski, Kalina Bontcheva

Named Entity Recognition (NER) is a key NLP task, which is all the more challenging on Web and user-generated content with their diverse and continuously changing language. This pa…

cs.CL2024

Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers

Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid +10

The increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the factual accuracy of their outputs. In this work, we…

cs.CL2016

Stance Detection with Bidirectional Conditional Encoding

Isabelle Augenstein, Tim Rocktäschel, Andreas Vlachos +1

Stance detection is the task of classifying the attitude expressed in a text towards a target such as Hillary Clinton to be "positive", negative" or "neutral". Previous work has as…

cs.CL2024

Can Transformers Learn -gram Language Models?

Anej Svete, Nadav Borenstein, Mike Zhou +2

Much theoretical work has described the ability of transformers to represent formal languages. However, linking theoretical results to empirical performance is not straightforward…

cs.CL2019

A Probabilistic Generative Model of Linguistic Typology

Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell +1

In the principles-and-parameters framework, the structural features of languages depend on parameters that may be toggled on or off, with a single parameter often dictating the sta…

cs.CL2026

Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthropomorphism, and Maxims

Siddhesh Milind Pawar, Sarah Masud, Haneul Yoo +2

Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses are communicated, not just…

cs.CL2026

Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025

Isabelle Augenstein

Language Models (LMs) acquire parametric knowledge from their training process, embedding it within their weights. The increasing scalability of LMs, however, poses significant cha…

cs.CL2025

Can Community Notes Replace Professional Fact-Checkers?

Nadav Borenstein, Greta Warren, Desmond Elliott +1

Two commonly employed strategies to combat the rise of misinformation on social media are (i) fact-checking by professional organisations and (ii) community moderation by platform…

cs.CL2019

Unsupervised Discovery of Gendered Language through Latent-Variable Modeling

Alexander Hoyle, Wolf-Sonkin, Hanna Wallach +2

Studying the ways in which language is gendered has long been an area of interest in sociolinguistics. Studies have explored, for example, the speech of male and female characters…

cs.CL2025

A Latent-Variable Model for Intrinsic Probing

Karolina Stańczak, Lucas Torroba Hennigen, Adina Williams +2

The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that…

cs.CL2021

Multi-Hop Fact Checking of Political Claims

Wojciech Ostrowski, Arnav Arora, Pepa Atanasova +1

Recent work has proposed multi-hop models and datasets for studying complex natural language reasoning. One notable task requiring multi-hop reasoning is fact checking, where a set…

cs.CL2020

2kenize: Tying Subword Sequences for Chinese Script Conversion

Pranav A, Isabelle Augenstein

Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP. Despite this, current approaches have poor performance because they do…

cs.CL2019

Combining Sentiment Lexica with a Multi-View Variational Autoencoder

Alexander Hoyle, Lawrence Wolf-Sonkin, Hanna Wallach +2

When assigning quantitative labels to a dataset, different methodologies may rely on different scales. In particular, when assigning polarities to words in a sentiment lexicon, ann…

cs.CL2024

DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models

Sara Vera Marjanović, Haeun Yu, Pepa Atanasova +3

Knowledge-intensive language understanding tasks require Language Models (LMs) to integrate relevant context, mitigating their inherent weaknesses, such as incomplete or outdated k…

cs.CL2024

SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages

Gayane Ghazaryan, Erik Arakelyan, Pasquale Minervini +1

Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other…

cs.CY2021

How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?

Indira Sen, Mattia Samory, Fabian Floeck +2

As NLP models are increasingly deployed in socially situated settings such as online abusive content detection, it is crucial to ensure that these models are robust. One way of imp…

cs.LG2024

Investigating the Impact of Model Instability on Explanations and Uncertainty

Sara Vera Marjanović, Isabelle Augenstein, Christina Lioma

Explainable AI methods facilitate the understanding of model behaviour, yet, small, imperceptible perturbations to inputs can vastly distort explanations. As these explanations are…

cs.CL2026

Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models

Haeun Yu, Seogyeong Jeong, Siddhesh Pawar +5

The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior wo…

cs.CL2023

Thorny Roses: Investigating the Dual Use Dilemma in Natural Language Processing

Lucie-Aimée Kaffee, Arnav Arora, Zeerak Talat +1

Dual use, the intentional, harmful reuse of technology and scientific artefacts, is a problem yet to be well-defined within the context of Natural Language Processing (NLP). Howeve…

cs.CL2017

A Supervised Approach to Extractive Summarisation of Scientific Papers

Ed Collins, Isabelle Augenstein, Sebastian Riedel

Automatic summarisation is a popular approach to reduce a document to its main arguments. Recent research in the area has focused on neural approaches to summarisation, which can b…

cs.CL2023

PHD: Pixel-Based Language Modeling of Historical Documents

Nadav Borenstein, Phillip Rust, Desmond Elliott +1

The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involve…

cs.CL2019

What do Language Representations Really Represent?

Johannes Bjerva, Robert Östling, Maria Han Veiga +2

A neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the cor…

cs.CL2021

A Primer on Contrastive Pretraining in Language Processing: Methods, Lessons Learned and Perspectives

Nils Rethmeier, Isabelle Augenstein

Modern natural language processing (NLP) methods employ self-supervised pretraining objectives such as masked language modeling to boost the performance of various application task…

cs.CL2022

Machine Reading, Fast and Slow: When Do Models "Understand" Language?

Sagnik Ray Choudhury, Anna Rogers, Isabelle Augenstein

Two of the most fundamental challenges in Natural Language Understanding (NLU) at present are: (a) how to establish whether deep learning-based models score highly on NLU benchmark…

cs.CL2024

Social Bias Probing: Fairness Benchmarking for Language Models

Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti +1

While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limitin…

cs.CL2023

Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection

Erik Arakelyan, Arnav Arora, Isabelle Augenstein

Stance Detection is concerned with identifying the attitudes expressed by an author towards a target of interest. This task spans a variety of domains ranging from social media opi…

cs.CL2019

Issue Framing in Online Discussion Fora

Mareike Hartmann, Tallulah Jansen, Isabelle Augenstein +1

In online discussion fora, speakers often make arguments for or against something, say birth control, by highlighting certain aspects of the topic. In social science, this is refer…

cs.CL2025

Probing Pre-Trained Language Models for Cross-Cultural Differences in Values

Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language…

cs.CL2025

Multi-Modal Framing Analysis of News

Arnav Arora, Srishti Yadav, Maria Antoniak +2

Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…

cs.CL2021

A Survey on Gender Bias in Natural Language Processing

Karolina Stanczak, Isabelle Augenstein

Language can be used as a means of reproducing and enforcing harmful stereotypes and biases and has been analysed as such in numerous research. In this paper, we present a survey o…

cs.CL2022

Generating Scientific Claims for Zero-Shot Scientific Fact Checking

Dustin Wright, David Wadden, Kyle Lo +4

Automated scientific fact checking is difficult due to the complexity of scientific language and a lack of significant amounts of training data, as annotation requires domain exper…

cs.CL2024

Understanding Fine-grained Distortions in Reports of Scientific Findings

Amelie Wührl, Dustin Wright, Roman Klinger +1

Distorted science communication harms individuals and society as it can lead to unhealthy behavior change and decrease trust in scientific institutions. Given the rapidly increasin…

cs.CL2017

Tracking Typological Traits of Uralic Languages in Distributed Language Representations

Johannes Bjerva, Isabelle Augenstein

Although linguistic typology has a long history, computational approaches have only recently gained popularity. The use of distributed representations in computational linguistics…

cs.CL2020

A Diagnostic Study of Explainability Techniques for Text Classification

Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma +1

Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales beh…

cs.CL2020

Generating Label Cohesive and Well-Formed Adversarial Claims

Pepa Atanasova, Dustin Wright, Isabelle Augenstein

Adversarial attacks reveal important vulnerabilities and flaws of trained models. One potent type of attack are universal adversarial triggers, which are individual n-grams that, w…

cs.CL2025

Unstructured Evidence Attribution for Long Context Query Focused Summarization

Dustin Wright, Zain Muhammad Mujahid, Lu Wang +2

Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the tru…

cs.LG2023

Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions

Lucie-Aimée Kaffee, Arnav Arora, Isabelle Augenstein

The moderation of content on online platforms is usually non-transparent. On Wikipedia, however, this discussion is carried out publicly and the editors are encouraged to use the c…

cs.HC2026

Show me the evidence: Evaluating the role of evidence and natural language explanations in AI-supported fact-checking

Greta Warren, Jingyi Sun, Irina Shklovski +1

Although much research has focused on AI explanations to support decisions in complex information-seeking tasks such as fact-checking, the role of evidence is surprisingly under-re…

cs.CL2026

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

Lucas Resck, Isabelle Augenstein, Anna Korhonen

Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptatio…

cs.CL2020

Claim Check-Worthiness Detection as Positive Unlabelled Learning

Dustin Wright, Isabelle Augenstein

As the first step of automatic fact checking, claim check-worthiness detection is a critical component of fact checking systems. There are multiple lines of research which study th…

cs.IR2021

University of Copenhagen Participation in TREC Health Misinformation Track 2020

Lucas Chaves Lima, Dustin Brandon Wright, Isabelle Augenstein +1

In this paper, we describe our participation in the TREC Health Misinformation Track 2020. We submitted runs to the Total Recall Task and 13 runs to the Ad Hoc task. Our appro…

cs.CL2025

Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift

Dustin Wright, Isabelle Augenstein

Selecting an effective training signal for machine learning tasks is difficult: expert annotations are expensive, and crowd-sourced annotations may not be reliable. Recent work has…

cs.CL2025

What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages

Nadav Borenstein, Anej Svete, Robin Chan +5

What can large language models learn? By definition, language models (LM) are distributions over strings. Therefore, an intuitive way of addressing the above question is to formali…

cs.CL2021

Time-Aware Evidence Ranking for Fact-Checking

Liesbeth Allein, Isabelle Augenstein, Marie-Francine Moens

Truth can vary over time. Fact-checking decisions on claim veracity should therefore take into account temporal information of both the claim and supporting or refuting evidence. I…

cs.CL2019

Domain Transfer in Dialogue Systems without Turn-Level Supervision

Joachim Bingel, Victor Petrén Bach Hansen, Ana Valeria Gonzalez +3

Task oriented dialogue systems rely heavily on specialized dialogue state tracking (DST) modules for dynamically predicting user intent throughout the conversation. State-of-the-ar…

cs.CL2019

X-WikiRE: A Large, Multilingual Resource for Relation Extraction as Machine Comprehension

Mostafa Abdou, Cezar Sas, Rahul Aralikatte +2

Although the vast majority of knowledge bases KBs are heavily biased towards English, Wikipedias do cover very different topics in different languages. Exploiting this, we introduc…

cs.CL2019

Back to the Future -- Sequential Alignment of Text Representations

Johannes Bjerva, Wouter Kouw, Isabelle Augenstein

Language evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural n…

cs.CV2025

With Great Backbones Comes Great Adversarial Transferability

Erik Arakelyan, Karen Hambardzumyan, Davit Papikyan +4

Advances in self-supervised learning (SSL) for machine vision have improved representation robustness and model performance, giving rise to pre-trained backbones like \emph{ResNet}…

stat.ML2018

Latent Multi-task Architecture Learning

Sebastian Ruder, Joachim Bingel, Isabelle Augenstein +1

Multi-task learning (MTL) allows deep neural networks to learn from related tasks by sharing parameters with other networks. In practice, however, MTL involves searching an enormou…

cs.CL2022

QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension

Anna Rogers, Matt Gardner, Isabelle Augenstein

Alongside huge volumes of research on deep learning models in NLP in the recent years, there has been also much work on benchmark datasets needed to track modeling progress. Questi…

cs.CL2020

Generating Fact Checking Explanations

Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma +1

Most existing work on automated fact checking is concerned with predicting the veracity of claims based on metadata, social network spread, language used in claims, and, more recen…

cs.CL2018

Nightmare at test time: How punctuation prevents parsers from generalizing

Anders Søgaard, Miryam de Lhoneux, Isabelle Augenstein

Punctuation is a strong indicator of syntactic structure, and parsers trained on text with punctuation often rely heavily on this signal. Punctuation is a diversion, however, since…

cs.CL2023

Explaining Interactions Between Text Spans

Sagnik Ray Choudhury, Pepa Atanasova, Isabelle Augenstein

Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehensi…

cs.CL2021

Generating Fluent Fact Checking Explanations with Unsupervised Post-Editing

Shailza Jolly, Pepa Atanasova, Isabelle Augenstein

Fact-checking systems have become important tools to verify fake and misguiding news. These systems become more trustworthy when human-readable explanations accompany the veracity…

cs.CL2023

Detecting Harmful Content On Online Platforms: What Platforms Need Vs. Where Research Efforts Go

Arnav Arora, Preslav Nakov, Momchil Hardalov +8

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and ha…