4 papers
TARDIS: Mitigating Temporal Misalignment via Representation Steering
Changho Shin, Xinya Yan, Suenggwan Jo +3
Language models often struggle with temporal misalignment, performance degradation caused by shifts in the temporal distribution of data. Continuously updating models to avoid degr…
Personalize Your LLM: Fake it then Align it
Yijing Zhang, Dyah Adila, Changho Shin +1
Personalizing large language models (LLMs) is essential for delivering tailored interactions that improve user experience. Many existing personalization methods require fine-tuning…
Weak-to-Strong Generalization Through the Data-Centric Lens
Changho Shin, John Cooper, Frederic Sala
The weak-to-strong generalization phenomenon is the driver for important machine learning applications including highly data-efficient learning and, most recently, performing super…
Is Free Self-Alignment Possible?
Dyah Adila, Changho Shin, Yijing Zhang +1
Aligning pretrained language models (LMs) often requires large-scale preference data and substantial computational resources. These costs become even more prohibitive for multi-obj…