11 papers
DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
Jason Lucas, Matt Murtagh, Ali Al-Lawati +3
Harmful content detectors, particularly disinformation classifiers, are predominantly developed and evaluated on Standard American English (SAE), leaving their robustness to dialec…
Topological Data Analysis Applications in Natural Language Processing: A Survey
Adaku Uchendu, Thai Le
The surge of data available on the Internet has driven the adoption of a wide range of computational methods for analyzing and extracting insights from large-scale data. Among thes…
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
Jason Lucas, Matt Murtagh-White, Adaku Uchendu +6
Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguis…
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
Dominik Macko, Aashish Anantha Ramakrishnan, Jason Samuel Lucas +4
Increased sophistication of large language models (LLMs) and the consequent quality of generated multilingual text raises concerns about potential disinformation misuse. While huma…
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
William Guo, Adaku Uchendu, Ana Smith
To mitigate the potential harms of Large Language Models (LLMs)generated text, researchers have proposed watermarking, a process of embedding detectable signals within text. With w…
Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification
Kenneth Alperin, Rohan Leekha, Adaku Uchendu +5
The increasing use of Artificial Intelligence (AI) technologies, such as Large Language Models (LLMs) has led to nontrivial improvements in various tasks, including accurate author…