3 papers
cs.CY2026
Questionnaire Responses Do not Capture the Safety of AI Agents
Max Hellrigel-Holderbaum, Edward James Young
As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such…
cs.CY2025
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
Leonard Dung, Max Hellrigel-Holderbaum
AGI Racing is the view that it is in the self-interest of major actors in AI development, especially powerful nations, to accelerate their frontier AI development to build highly c…
cs.CY2025
Misalignment or misuse? The AGI alignment tradeoff
Max Hellrigel-Holderbaum, Leonard Dung
Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We…