Publications (151)
What Can We Do to Improve Peer Review in NLP?
Anna Rogers, Isabelle Augenstein
Peer review is our best tool for judging the quality of conference submissions, but it is becoming increasingly spurious. We argue that a part of the problem is that the reviewers…
Multi-task Learning of Pairwise Sequence Classification Tasks Over Disparate Label Spaces
Isabelle Augenstein, Sebastian Ruder, Anders Søgaard
We combine multi-task learning and semi-supervised learning by inducing a joint embedding space between disparate label spaces and learning transfer functions between label embeddi…
Disembodied Machine Learning: On the Illusion of Objectivity in NLP
Zeerak Waseem, Smarika Lulz, Joachim Bingel +1
Machine Learning seeks to identify and encode bodies of knowledge within provided datasets. However, data encodes subjective content, which determines the possible outcomes of the…
Copenhagen at CoNLL--SIGMORPHON 2018: Multilingual Inflection in Context with Explicit Morphosyntactic Decoding
Yova Kementchedjhieva, Johannes Bjerva, Isabelle Augenstein
This paper documents the Team Copenhagen system which placed first in the CoNLL--SIGMORPHON 2018 shared task on universal morphological reinflection, Task 2 with an overall accurac…
A simple but tough-to-beat baseline for the Fake News Challenge stance detection task
Benjamin Riedel, Isabelle Augenstein, Georgios P. Spithourakis +1
Identifying public misinformation is a complicated and challenging task. An important part of checking the veracity of a specific claim is to evaluate the stance different news sou…
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection
Indira Sen, Mattia Samory, Claudia Wagner +1
Counterfactually Augmented Data (CAD) aims to improve out-of-domain generalizability, an indicator of model robustness. The improvement is credited with promoting core features of…
Multi-Task Learning of Keyphrase Boundary Classification
Isabelle Augenstein, Anders Søgaard
Keyphrase boundary classification (KBC) is the task of detecting keyphrases in scientific articles and labelling them with respect to predefined types. Although important in practi…
BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
Sekh Mainul Islam, Nadav Borenstein, Siddhesh Milind Pawar +3
Understanding biases and stereotypes encoded in the weights of Large Language Models (LLMs) is crucial for developing effective mitigation strategies. However, biased behaviour is…
Does Typological Blinding Impede Cross-Lingual Sharing?
Johannes Bjerva, Isabelle Augenstein
Bridging the performance gap between high- and low-resource languages has been the focus of much previous work. Typological features from databases such as the World Atlas of Langu…
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns
Amalie Brogaard Pauli, Maria Barrett, Max Müller-Eberstein +2
Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to influence and persuade. Prior work has sh…
Interpretable Debiasing of Vision-Language Models for Social Fairness
Na Min An, Yoonna Jang, Yusuke Hirota +3
The rapid advancement of Vision-Language models (VLMs) has raised growing concerns that their black-box reasoning processes could lead to unintended forms of social bias. Current d…
CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding
Dustin Wright, Isabelle Augenstein
Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annota…
Output Vector Editing for Memorization Mitigation in Large Language Models
Ahmad Dawar Hakimi, Kaiwei Lei, Isabelle Augenstein +1
Large language models memorize and reproduce sequences from their training data, creating privacy, copyright, and security risks. Existing neuron-level mitigation methods equate ed…
Investigating Human Values in Online Communities
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee +1
Studying human values is instrumental for cross-cultural research, enabling a better understanding of preferences and behaviour of society at large and communities therein. To stud…
Numerically Grounded Language Models for Semantic Error Correction
Georgios P. Spithourakis, Isabelle Augenstein, Sebastian Riedel
Semantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction. Current approaches generally fo…
TX-Ray: Quantifying and Explaining Model-Knowledge Transfer in (Un-)Supervised NLP
Nils Rethmeier, Vageesh Kumar Saxena, Isabelle Augenstein
While state-of-the-art NLP explainability (XAI) methods focus on explaining per-sample decisions in supervised end or probing tasks, this is insufficient to explain and quantify mo…
Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models
Karolina StaÅczak, Edoardo Ponti, Lucas Torroba Hennigen +2
The success of multilingual pre-trained models is underpinned by their ability to learn representations shared by multiple languages even in absence of any explicit supervision. Ho…
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings
Johannes Bjerva, Isabelle Augenstein
A core part of linguistic typology is the classification of languages according to linguistic properties, such as those detailed in the World Atlas of Language Structure (WALS). Do…
Epistemic Diversity and Knowledge Collapse in Large Language Models
Dustin Wright, Sarah Masud, Jared Moore +5
Large language models (LLMs) tend to generate homogenous texts, which may impact the diversity of knowledge generated across different outputs. Given their potential to replace exi…
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
Indira Sen, Dennis Assenmacher, Mattia Samory +3
NLP models are used in a variety of critical social computing tasks, such as detecting sexist, racist, or otherwise hateful content. Therefore, it is imperative that these models a…
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
Lovisa Hagström, Sara Vera MarjanoviÄ, Haeun Yu +5
Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can…
Community Moderation and the New Epistemology of Fact Checking on Social Media
Isabelle Augenstein, Michiel Bakker, Tanmoy Chakraborty +13
Social media platforms have traditionally relied on internal moderation teams and partnerships with independent fact-checking organizations to identify and flag misleading content.…
Inducing Language-Agnostic Multilingual Representations
Wei Zhao, Steffen Eger, Johannes Bjerva +1
Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world. However, they currently require large pretraining…
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
Amalie Brogaard Pauli, Isabelle Augenstein, Ira Assent
Large language models (LLMs) make it easy to rewrite a text in any style -- e.g. to make it more polite, persuasive, or more positive -- but evaluation thereof is not straightforwa…
Joint Emotion Label Space Modelling for Affect Lexica
Luna De Bruyne, Pepa Atanasova, Isabelle Augenstein
Emotion lexica are commonly used resources to combat data poverty in automatic emotion detection. However, vocabulary coverage issues, differences in construction method and discre…
SIGTYP 2020 Shared Task: Prediction of Typological Features
Johannes Bjerva, Elizabeth Salesky, Sabrina J. Mielke +6
Typological knowledge bases (KBs) such as WALS (Dryer and Haspelmath, 2013) contain information about linguistic properties of the world's languages. They have been shown to be use…
Zero-Shot Cross-Lingual Transfer with Meta Learning
Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva +1
Learning what to share between tasks has been a topic of great importance recently, as strategic sharing of knowledge has been shown to improve downstream task performance. This is…
A Neighbourhood Framework for Resource-Lean Content Flagging
Sheikh Muhammad Sarwar, Dimitrina Zlatkova, Momchil Hardalov +3
We propose a novel framework for cross-lingual content flagging with limited target-language data, which significantly outperforms prior work in terms of predictive performance. Th…
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein
Natural Language Explanations (NLEs) describe how Large Language Models (LLMs) make decisions by drawing on external Context Knowledge (CK) and Parametric Knowledge (PK). Understan…
Semi-Supervised Exaggeration Detection of Health Science Press Releases
Dustin Wright, Isabelle Augenstein
Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scienti…
Cross-Domain Label-Adaptive Stance Detection
Momchil Hardalov, Arnav Arora, Preslav Nakov +1
Stance detection concerns the classification of a writer's viewpoint towards a target. There are different task variants, e.g., stance of a tweet vs. a full article, or stance with…
Unsupervised Evaluation for Question Answering with Transformers
Lukas Muttenthaler, Isabelle Augenstein, Johannes Bjerva
It is challenging to automatically evaluate the answer of a QA model at inference time. Although many models provide confidence scores, and simple heuristics can go a long way towa…
Faithfulness Tests for Natural Language Explanations
Pepa Atanasova, Oana-Maria Camburu, Christina Lioma +3
Explanations of neural models aim to reveal a model's decision-making process for its predictions. However, recent work shows that current methods giving explanations such as salie…
Graph-Guided Textual Explanation Generation Framework
Shuzhou Yuan, Jingyi Sun, Ran Zhang +4
Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions. However, recent work has questioned…
Revealing Fine-Grained Values and Opinions in Large Language Models
Dustin Wright, Arnav Arora, Nadav Borenstein +3
Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…
The Causal Influence of Grammatical Gender on Distributional Semantics
Karolina StaÅczak, Kevin Du, Adina Williams +2
How much meaning influences gender assignment across languages is an active area of research in linguistics and cognitive science. We can view current approaches as aiming to deter…
Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings
Malte Ostendorff, Nils Rethmeier, Isabelle Augenstein +2
Learning scientific document representations can be substantially improved through contrastive learning objectives, where the challenge lies in creating positive and negative train…
Discourse-Aware Rumour Stance Classification in Social Media Using Sequential Classifiers
Arkaitz Zubiaga, Elena Kochkina, Maria Liakata +5
Rumour stance classification, defined as classifying the stance of specific social media posts into one of supporting, denying, querying or commenting on an earlier post, is becomi…
Transductive Auxiliary Task Self-Training for Neural Multi-Task Models
Johannes Bjerva, Katharina Kann, Isabelle Augenstein
Multi-task learning and self-training are two common ways to improve a machine learning model's performance in settings with limited training data. Drawing heavily on ideas from th…
Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
Greta Warren, Irina Shklovski, Isabelle Augenstein
The pervasiveness of large language models and generative AI in online media has amplified the need for effective automated fact-checking to assist fact-checkers in tackling the in…
Multi-Sense Language Modelling
Andrea Lekkas, Peter Schneider-Kamp, Isabelle Augenstein
The effectiveness of a language model is influenced by its token representations, which must encode contextual information and handle the same word form having a plurality of meani…
Uncovering Probabilistic Implications in Typological Knowledge Bases
Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell +1
The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have post…
Retrieval-based Goal-Oriented Dialogue Generation
Ana Valeria Gonzalez, Isabelle Augenstein, Anders Søgaard
Most research on dialogue has focused either on dialogue generation for openended chit chat or on state tracking for goal-directed dialogue. In this work, we explore a hybrid appro…
Quantifying Gender Biases Towards Politicians on Reddit
Sara Marjanovic, Karolina StaÅczak, Isabelle Augenstein
Despite attempts to increase gender parity in politics, global efforts have struggled to ensure equal female representation. This is likely tied to implicit gender biases against w…
Modeling Information Change in Science Communication with Semantically Matched Paraphrases
Dustin Wright, Jiaxin Pei, David Jurgens +1
Whether the media faithfully communicate scientific information has long been a core issue to the science community. Automatically identifying paraphrased scientific findings could…
Generalisation in Named Entity Recognition: A Quantitative Analysis
Isabelle Augenstein, Leon Derczynski, Kalina Bontcheva
Named Entity Recognition (NER) is a key NLP task, which is all the more challenging on Web and user-generated content with their diverse and continuously changing language. This pa…
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid +10
The increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the factual accuracy of their outputs. In this work, we…
Stance Detection with Bidirectional Conditional Encoding
Isabelle Augenstein, Tim Rocktäschel, Andreas Vlachos +1
Stance detection is the task of classifying the attitude expressed in a text towards a target such as Hillary Clinton to be "positive", negative" or "neutral". Previous work has as…
Can Transformers Learn -gram Language Models?
Anej Svete, Nadav Borenstein, Mike Zhou +2
Much theoretical work has described the ability of transformers to represent formal languages. However, linking theoretical results to empirical performance is not straightforward…
A Probabilistic Generative Model of Linguistic Typology
Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell +1
In the principles-and-parameters framework, the structural features of languages depend on parameters that may be toggled on or off, with a single parameter often dictating the sta…
Not What, But How: A Framework for Auditing LLM Responses across Positioning, Generalization, Anthropomorphism, and Maxims
Siddhesh Milind Pawar, Sarah Masud, Haneul Yoo +2
Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses are communicated, not just…
Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025
Isabelle Augenstein
Language Models (LMs) acquire parametric knowledge from their training process, embedding it within their weights. The increasing scalability of LMs, however, poses significant cha…
Can Community Notes Replace Professional Fact-Checkers?
Nadav Borenstein, Greta Warren, Desmond Elliott +1
Two commonly employed strategies to combat the rise of misinformation on social media are (i) fact-checking by professional organisations and (ii) community moderation by platform…
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling
Alexander Hoyle, Wolf-Sonkin, Hanna Wallach +2
Studying the ways in which language is gendered has long been an area of interest in sociolinguistics. Studies have explored, for example, the speech of male and female characters…
A Latent-Variable Model for Intrinsic Probing
Karolina StaÅczak, Lucas Torroba Hennigen, Adina Williams +2
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that…
Multi-Hop Fact Checking of Political Claims
Wojciech Ostrowski, Arnav Arora, Pepa Atanasova +1
Recent work has proposed multi-hop models and datasets for studying complex natural language reasoning. One notable task requiring multi-hop reasoning is fact checking, where a set…
2kenize: Tying Subword Sequences for Chinese Script Conversion
Pranav A, Isabelle Augenstein
Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP. Despite this, current approaches have poor performance because they do…
Combining Sentiment Lexica with a Multi-View Variational Autoencoder
Alexander Hoyle, Lawrence Wolf-Sonkin, Hanna Wallach +2
When assigning quantitative labels to a dataset, different methodologies may rely on different scales. In particular, when assigning polarities to words in a sentiment lexicon, ann…
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
Sara Vera MarjanoviÄ, Haeun Yu, Pepa Atanasova +3
Knowledge-intensive language understanding tasks require Language Models (LMs) to integrate relevant context, mitigating their inherent weaknesses, such as incomplete or outdated k…
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
Gayane Ghazaryan, Erik Arakelyan, Pasquale Minervini +1
Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other…
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?
Indira Sen, Mattia Samory, Fabian Floeck +2
As NLP models are increasingly deployed in socially situated settings such as online abusive content detection, it is crucial to ensure that these models are robust. One way of imp…
Investigating the Impact of Model Instability on Explanations and Uncertainty
Sara Vera MarjanoviÄ, Isabelle Augenstein, Christina Lioma
Explainable AI methods facilitate the understanding of model behaviour, yet, small, imperceptible perturbations to inputs can vastly distort explanations. As these explanations are…
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
Haeun Yu, Seogyeong Jeong, Siddhesh Pawar +5
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of LLMs' representations of different cultures. Prior wo…
Thorny Roses: Investigating the Dual Use Dilemma in Natural Language Processing
Lucie-Aimée Kaffee, Arnav Arora, Zeerak Talat +1
Dual use, the intentional, harmful reuse of technology and scientific artefacts, is a problem yet to be well-defined within the context of Natural Language Processing (NLP). Howeve…
A Supervised Approach to Extractive Summarisation of Scientific Papers
Ed Collins, Isabelle Augenstein, Sebastian Riedel
Automatic summarisation is a popular approach to reduce a document to its main arguments. Recent research in the area has focused on neural approaches to summarisation, which can b…
PHD: Pixel-Based Language Modeling of Historical Documents
Nadav Borenstein, Phillip Rust, Desmond Elliott +1
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involve…
What do Language Representations Really Represent?
Johannes Bjerva, Robert Ãstling, Maria Han Veiga +2
A neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the cor…
A Primer on Contrastive Pretraining in Language Processing: Methods, Lessons Learned and Perspectives
Nils Rethmeier, Isabelle Augenstein
Modern natural language processing (NLP) methods employ self-supervised pretraining objectives such as masked language modeling to boost the performance of various application task…
Machine Reading, Fast and Slow: When Do Models "Understand" Language?
Sagnik Ray Choudhury, Anna Rogers, Isabelle Augenstein
Two of the most fundamental challenges in Natural Language Understanding (NLU) at present are: (a) how to establish whether deep learning-based models score highly on NLU benchmark…
Social Bias Probing: Fairness Benchmarking for Language Models
Marta Marchiori Manerba, Karolina StaÅczak, Riccardo Guidotti +1
While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limitin…
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection
Erik Arakelyan, Arnav Arora, Isabelle Augenstein
Stance Detection is concerned with identifying the attitudes expressed by an author towards a target of interest. This task spans a variety of domains ranging from social media opi…
Issue Framing in Online Discussion Fora
Mareike Hartmann, Tallulah Jansen, Isabelle Augenstein +1
In online discussion fora, speakers often make arguments for or against something, say birth control, by highlighting certain aspects of the topic. In social science, this is refer…
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein
Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language…
Multi-Modal Framing Analysis of News
Arnav Arora, Srishti Yadav, Maria Antoniak +2
Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…
A Survey on Gender Bias in Natural Language Processing
Karolina Stanczak, Isabelle Augenstein
Language can be used as a means of reproducing and enforcing harmful stereotypes and biases and has been analysed as such in numerous research. In this paper, we present a survey o…
Generating Scientific Claims for Zero-Shot Scientific Fact Checking
Dustin Wright, David Wadden, Kyle Lo +4
Automated scientific fact checking is difficult due to the complexity of scientific language and a lack of significant amounts of training data, as annotation requires domain exper…
Understanding Fine-grained Distortions in Reports of Scientific Findings
Amelie Wührl, Dustin Wright, Roman Klinger +1
Distorted science communication harms individuals and society as it can lead to unhealthy behavior change and decrease trust in scientific institutions. Given the rapidly increasin…
Tracking Typological Traits of Uralic Languages in Distributed Language Representations
Johannes Bjerva, Isabelle Augenstein
Although linguistic typology has a long history, computational approaches have only recently gained popularity. The use of distributed representations in computational linguistics…
A Diagnostic Study of Explainability Techniques for Text Classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma +1
Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity. Efforts to make the rationales beh…
Generating Label Cohesive and Well-Formed Adversarial Claims
Pepa Atanasova, Dustin Wright, Isabelle Augenstein
Adversarial attacks reveal important vulnerabilities and flaws of trained models. One potent type of attack are universal adversarial triggers, which are individual n-grams that, w…
Unstructured Evidence Attribution for Long Context Query Focused Summarization
Dustin Wright, Zain Muhammad Mujahid, Lu Wang +2
Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the tru…
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
Lucie-Aimée Kaffee, Arnav Arora, Isabelle Augenstein
The moderation of content on online platforms is usually non-transparent. On Wikipedia, however, this discussion is carried out publicly and the editors are encouraged to use the c…
Show me the evidence: Evaluating the role of evidence and natural language explanations in AI-supported fact-checking
Greta Warren, Jingyi Sun, Irina Shklovski +1
Although much research has focused on AI explanations to support decisions in complex information-seeking tasks such as fact-checking, the role of evidence is surprisingly under-re…
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
Lucas Resck, Isabelle Augenstein, Anna Korhonen
Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptatio…
Claim Check-Worthiness Detection as Positive Unlabelled Learning
Dustin Wright, Isabelle Augenstein
As the first step of automatic fact checking, claim check-worthiness detection is a critical component of fact checking systems. There are multiple lines of research which study th…
University of Copenhagen Participation in TREC Health Misinformation Track 2020
Lucas Chaves Lima, Dustin Brandon Wright, Isabelle Augenstein +1
In this paper, we describe our participation in the TREC Health Misinformation Track 2020. We submitted runs to the Total Recall Task and 13 runs to the Ad Hoc task. Our appro…
Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
Dustin Wright, Isabelle Augenstein
Selecting an effective training signal for machine learning tasks is difficult: expert annotations are expensive, and crowd-sourced annotations may not be reliable. Recent work has…
What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
Nadav Borenstein, Anej Svete, Robin Chan +5
What can large language models learn? By definition, language models (LM) are distributions over strings. Therefore, an intuitive way of addressing the above question is to formali…
Time-Aware Evidence Ranking for Fact-Checking
Liesbeth Allein, Isabelle Augenstein, Marie-Francine Moens
Truth can vary over time. Fact-checking decisions on claim veracity should therefore take into account temporal information of both the claim and supporting or refuting evidence. I…
Domain Transfer in Dialogue Systems without Turn-Level Supervision
Joachim Bingel, Victor Petrén Bach Hansen, Ana Valeria Gonzalez +3
Task oriented dialogue systems rely heavily on specialized dialogue state tracking (DST) modules for dynamically predicting user intent throughout the conversation. State-of-the-ar…
X-WikiRE: A Large, Multilingual Resource for Relation Extraction as Machine Comprehension
Mostafa Abdou, Cezar Sas, Rahul Aralikatte +2
Although the vast majority of knowledge bases KBs are heavily biased towards English, Wikipedias do cover very different topics in different languages. Exploiting this, we introduc…
Back to the Future -- Sequential Alignment of Text Representations
Johannes Bjerva, Wouter Kouw, Isabelle Augenstein
Language evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural n…
With Great Backbones Comes Great Adversarial Transferability
Erik Arakelyan, Karen Hambardzumyan, Davit Papikyan +4
Advances in self-supervised learning (SSL) for machine vision have improved representation robustness and model performance, giving rise to pre-trained backbones like \emph{ResNet}…
Latent Multi-task Architecture Learning
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein +1
Multi-task learning (MTL) allows deep neural networks to learn from related tasks by sharing parameters with other networks. In practice, however, MTL involves searching an enormou…
QA Dataset Explosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension
Anna Rogers, Matt Gardner, Isabelle Augenstein
Alongside huge volumes of research on deep learning models in NLP in the recent years, there has been also much work on benchmark datasets needed to track modeling progress. Questi…
Generating Fact Checking Explanations
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma +1
Most existing work on automated fact checking is concerned with predicting the veracity of claims based on metadata, social network spread, language used in claims, and, more recen…
Nightmare at test time: How punctuation prevents parsers from generalizing
Anders Søgaard, Miryam de Lhoneux, Isabelle Augenstein
Punctuation is a strong indicator of syntactic structure, and parsers trained on text with punctuation often rely heavily on this signal. Punctuation is a diversion, however, since…
Explaining Interactions Between Text Spans
Sagnik Ray Choudhury, Pepa Atanasova, Isabelle Augenstein
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehensi…
Generating Fluent Fact Checking Explanations with Unsupervised Post-Editing
Shailza Jolly, Pepa Atanasova, Isabelle Augenstein
Fact-checking systems have become important tools to verify fake and misguiding news. These systems become more trustworthy when human-readable explanations accompany the veracity…
Detecting Harmful Content On Online Platforms: What Platforms Need Vs. Where Research Efforts Go
Arnav Arora, Preslav Nakov, Momchil Hardalov +8
The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and ha…