3 papers
cs.LG2026
Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs
Elizaveta Tennant, Benjamin Henke, Anita Keshmirian +5
As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional eval…
cs.AI2026
Evaluating Language Models for Harmful Manipulation
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9
Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…
cs.AI2024
STAR: SocioTechnical Approach to Red Teaming Language Models
Laura Weidinger, John Mellor, Bernat Guillen Pegueroles +9
This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions:…