1 paper
Ioannis Stylianou, Sven Ewan Shepstone, Jon Francombe +2
Large Language Model alignment typically relies on learned proxy reward models, which significantly increase the memory footprint during training and are notoriously prone to insta…