From the 1 of 1 linked paper with an AI index.
1 paper · 1 filter
Desiree Cho, Cameron Tice, Bernie Hogan +4
The paper investigates inserting constitutionally‑derived content during midtraining of large language models to improve the durability of alignment, showing reduced blackmail tend…