3 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CY2025★ 3 cited
Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts
Severin Field
The development of artificial general intelligence (AGI) is likely to be one of humanity's most consequential technological advancements. Leading AI labs and scientists have called…
cs.LG2024
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
Anthony Costarelli, Mat Allen, Severin Field
As Large Language Models (LLMs) become increasingly integrated into our daily lives, the potential harms from deceptive behavior underlie the need for faithfully interpreting their…
cs.CR2024
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
Nathalie Kirch, Constantin Weisser, Severin Field +2
Jailbreaks have been a central focus of research regarding the safety and reliability of large language models (LLMs), yet the mechanisms underlying these attacks remain poorly und…