3 papers
cs.LG2026
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin +1
Alignment audits aim to robustly identify hidden goals from strategic, situationally aware misaligned models. Despite this threat model, existing auditing methods have not been sys…
cs.LG2025
Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework
Yannick Metz, David Lindner, Raphaël Baur +1
Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…
cs.AI2024
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
Kai Fronsdal, David Lindner
We propose a suite of tasks to evaluate the instrumental self-reasoning ability of large language model (LLM) agents. Instrumental self-reasoning ability could improve adaptability…