36 citations · 64 across the 19 of their papers we have counts for
Showing 2026 · cs.CLShow all
2 papers · 2 filters
cs.CL2026
Model Unlearning Objectives Vary for Distinct Language Functions
Berk Atil, Vipul Gupta, Rebecca J. Passonneau
Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-training uses different objectiv…
cs.CL2026
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
Berk Atil, Rebecca J. Passonneau, Ninareh Mehrabi
Toxicity detection is inherently subjective, shaped by the diverse perspectives and social priors of different demographic groups. While ``pluralistic'' modeling as used in economi…