From the 1 of 9 linked papers with an AI index.
1 paper · 1 filter
Desiree Cho, Cameron Tice, Bernie Hogan +4
The paper investigates inserting constitutionally‑derived content during midtraining of large language models to improve the durability of alignment, showing reduced blackmail tend…