1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CYShow all
2 papers · 1 filter
cs.CY2025
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
Yuan Yuan, Tina Sriskandarajah, Anna-Luisa Brakman +4
Large Language Models used in ChatGPT have traditionally been trained to learn a refusal boundary: depending on the user's intent, the model is taught to either fully comply or out…
cs.CY2024
First-Person Fairness in Chatbots
Tyna Eloundou, Alex Beutel, David G. Robinson +7
Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-maki…