2 papers
cs.CL2026
Abliteration Mitigation via Refusal Aliases
Nathan Truong
Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent…
cs.AI2026
LLM Scheming Inversely Scales with Pretraining Language Coverage
Nathan Truong, Aryan Panda, Rayming Ye +2
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-con…