Publications (22)
Epistemic Diversity and Knowledge Collapse in Large Language Models
Dustin Wright, Sarah Masud, Jared Moore +5
Large language models (LLMs) tend to generate homogenous texts, which may impact the diversity of knowledge generated across different outputs. Given their potential to replace exi…
BMRS: Bayesian Model Reduction for Structured Pruning
Dustin Wright, Christian Igel, Raghavendra Selvan
Modern neural networks are often massively overparameterized leading to high compute costs during training and at inference. One effective method to improve both the compute and en…
Stress Testing Factual Consistency Metrics for Long-Document Summarization
Zain Muhammad Mujahid, Dustin Wright, Isabelle Augenstein
Evaluating the factual consistency of abstractive text summarization remains a significant challenge, particularly for long documents, where conventional metrics struggle with inpu…
Longitudinal Citation Prediction using Temporal Graph Neural Networks
Andreas Nugaard Holm, Barbara Plank, Dustin Wright +1
Citation count prediction is the task of predicting the number of citations a paper has gained after a period of time. Prior work viewed this as a static prediction task. As papers…
Modeling Information Change in Science Communication with Semantically Matched Paraphrases
Dustin Wright, Jiaxin Pei, David Jurgens +1
Whether the media faithfully communicate scientific information has long been a core issue to the science community. Automatically identifying paraphrased scientific findings could…
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
Kevin Roitero, Dustin Wright, Michael Soprano +2
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessment…
Transformer Based Multi-Source Domain Adaptation
Dustin Wright, Isabelle Augenstein
In practical machine learning settings, the data on which a model must make predictions often come from a different distribution than the data it was trained on. Here, we investiga…
Generating Scientific Claims for Zero-Shot Scientific Fact Checking
Dustin Wright, David Wadden, Kyle Lo +4
Automated scientific fact checking is difficult due to the complexity of scientific language and a lack of significant amounts of training data, as annotation requires domain exper…
Understanding Fine-grained Distortions in Reports of Scientific Findings
Amelie Wührl, Dustin Wright, Roman Klinger +1
Distorted science communication harms individuals and society as it can lead to unhealthy behavior change and decrease trust in scientific institutions. Given the rapidly increasin…
Machine Understanding of Scientific Language
Dustin Wright
Scientific information expresses human understanding of nature. This knowledge is largely disseminated in different forms of text, including scientific papers, news articles, and d…
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
Jonathan Ivey, Shivani Kumar, Jiayu Liu +12
Studying and building datasets for dialogue tasks is both expensive and time-consuming due to the need to recruit, train, and collect data from study participants. In response, muc…
Generating Label Cohesive and Well-Formed Adversarial Claims
Pepa Atanasova, Dustin Wright, Isabelle Augenstein
Adversarial attacks reveal important vulnerabilities and flaws of trained models. One potent type of attack are universal adversarial triggers, which are individual n-grams that, w…
Semi-Supervised Exaggeration Detection of Health Science Press Releases
Dustin Wright, Isabelle Augenstein
Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scienti…
Unstructured Evidence Attribution for Long Context Query Focused Summarization
Dustin Wright, Zain Muhammad Mujahid, Lu Wang +2
Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the tru…
CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding
Dustin Wright, Isabelle Augenstein
Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annota…
Modeling Public Perceptions of Science in Media
Jiaxin Pei, Dustin Wright, Isabelle Augenstein +1
Effectively engaging the public with science is vital for fostering trust and understanding in our scientific community. Yet, with an ever-growing volume of information, science co…
Claim Check-Worthiness Detection as Positive Unlabelled Learning
Dustin Wright, Isabelle Augenstein
As the first step of automatic fact checking, claim check-worthiness detection is a critical component of fact checking systems. There are multiple lines of research which study th…
Revisiting Softmax for Uncertainty Approximation in Text Classification
Andreas Nugaard Holm, Dustin Wright, Isabelle Augenstein
Uncertainty approximation in text classification is an important area with applications in domain adaptation and interpretability. One of the most widely used uncertainty approxima…
Rethinking Recurrent Latent Variable Model for Music Composition
Eunjeong Stella Koh, Shlomo Dubnov, Dustin Wright
We present a model for capturing musical features and creating novel sequences of music, called the Convolutional Variational Recurrent Neural Network. To generate sequential data,…
Efficiency is Not Enough: A Critical Perspective of Environmentally Sustainable AI
Dustin Wright, Christian Igel, Gabrielle Samuel +1
Artificial intelligence (AI) is currently spearheaded by machine learning (ML) methods such as deep learning which have accelerated progress on many tasks thought to be out of reac…
Revealing Fine-Grained Values and Opinions in Large Language Models
Dustin Wright, Arnav Arora, Nadav Borenstein +3
Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…
Aggregating Soft Labels from Crowd Annotations Improves Uncertainty Estimation Under Distribution Shift
Dustin Wright, Isabelle Augenstein
Selecting an effective training signal for machine learning tasks is difficult: expert annotations are expensive, and crowd-sourced annotations may not be reliable. Recent work has…