Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Peer-Preservation in Frontier Models
Yujin Potter, Nicholas Crispino, Vincent Siu +2
Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can also exhibit misaligned behaviors in def…
cs.CL2025
COSMIC: Generalized Refusal Direction Identification in LLM Activations
Vincent Siu, Nicholas Crispino, Zihao Yu +5
Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often…