2 papers
cs.CL2026
Constitutional Midtraining: Content Presence Drives Alignment Gains
Desiree Cho, Cameron Tice, Bernie Hogan +4
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when c…
cs.CY2024
Towards a Harms Taxonomy of AI Likeness Generation
Ben Bariach, Bernie Hogan, Keegan McBride
Generative artificial intelligence models, when trained on a sufficient number of a person's images, can replicate their identifying features in a photorealistic manner. We refer t…