90 citations · 137 across the 12 of their papers we have counts for
6 papers · 1 filter
Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas
Louis Kwok, Michal Bravansky, Lewis D. Griffin
The success of Large Language Models (LLMs) in multicultural environments hinges on their ability to understand users' diverse cultural backgrounds. We measure this capability by h…
Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
Maximilian Mozes, Xuanli He, Bennett Kleinberg +1
Spurred by the recent rapid increase in the development and distribution of large language models (LLMs) across industry and academia, much recent work has drawn attention to safet…
Susceptibility to Influence of Large Language Models
Lewis D Griffin, Bennett Kleinberg, Maximilian Mozes +4
Two studies tested the hypothesis that a Large Language Model (LLM) can be used to model psychological change following exposure to influential input. The first study tested a gene…
Identifying Human Strategies for Generating Word-Level Adversarial Examples
Maximilian Mozes, Bennett Kleinberg, Lewis D. Griffin
Adversarial examples in NLP are receiving increasing research attention. One line of investigation is the generation of word-level adversarial examples against fine-tuned Transform…
Self-Supervised Losses for One-Class Textual Anomaly Detection
Kimberly T. Mai, Toby Davies, Lewis D. Griffin
Current deep learning methods for anomaly detection in text rely on supervisory signals in inliers that may be unobtainable or bespoke architectures that are difficult to tune. We…
Contrasting Human- and Machine-Generated Word-Level Adversarial Examples for Text Classification
Maximilian Mozes, Max Bartolo, Pontus Stenetorp +2
Research shows that natural language processing models are generally considered to be vulnerable to adversarial attacks; but recent work has drawn attention to the issue of validat…