collaborators

8 papers

cs.CL2026

Model Unlearning Objectives Vary for Distinct Language Functions

Berk Atil, Vipul Gupta, Rebecca J. Passonneau

Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-training uses different objectiv…

cs.AI2026

Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning

Jihyun Janice Ahn, Ryo Kamoi, Berk Atil +34

LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify…

cs.CL2026

Self-Explaining Hate Speech Detection with Moral Rationales

Francielle Vargas, Jackson Trager, Diego Alves +6

Existing hate speech detection models are often opaque and rely on surface-level lexical cues, which makes them vulnerable to spurious correlations and limits robustness, interpret…

cs.CL2026

Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling

Berk Atil, Rebecca J. Passonneau, Ninareh Mehrabi

Toxicity detection is inherently subjective, shaped by the diverse perspectives and social priors of different demographic groups. While ``pluralistic'' modeling as used in economi…

cs.CL2026

Something Just Like TRuST : Toxicity Recognition of Span and Target

Berk Atil, Namrata Sureddy, Rebecca J. Passonneau

Toxic language includes content that is offensive, abusive, or that promotes harm. Progress in preventing toxic output from large language models (LLMs) is hampered by inconsistent…

cs.CL2025

Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?

Berk Atil, Rebecca J. Passonneau, Fred Morstatter

Large language models (LLMs) undergo safety alignment after training and tuning, yet recent work shows that safety can be bypassed through jailbreak attacks. While many jailbreaks…