Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment
Jason R. Brown, Patrick Leask, Lev McKinney
Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned beh…
cs.AI2026
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff +18
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning intern…