20 citations · 28 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Evaluating Language Models for Harmful Manipulation
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9
Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…
cs.AI2024
STAR: SocioTechnical Approach to Red Teaming Language Models
Laura Weidinger, John Mellor, Bernat Guillen Pegueroles +9
This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions:…