13 papers
Evaluating Large Language Models for Antisemitic Incident Classification
Karina Halevy, Julia Mendelsohn, Chan Young Park +2
Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We…
Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues
Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov
As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behaviors are genuine or superficia…
Scaling Participation in Modular AI Systems
Shangbin Feng, Yike Wang, Weijia Shi +3
Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by the few -- a centralized mark…
MoCo: A One-Stop Shop for Model Collaboration Research
Shangbin Feng, Yuyang Bai, Ziyuan Yang +17
Advancing beyond single monolithic language models (LMs), recent research increasingly recognizes the importance of model collaboration, where multiple LMs collaborate, compose, an…
Biased AI can Influence Political Decision-Making
Jillian Fisher, Shangbin Feng, Robert Aron +6
As modern large language models (LLMs) become integral to everyday tasks, concerns about their inherent biases and their potential impact on human decision-making have emerged. Whi…
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
Rui Xin, Niloofar Mireshghallah, Shuyue Stella Li +6
Sanitizing sensitive text data typically involves removing personally identifiable information (PII) or generating synthetic data under the assumption that these methods adequately…