4 papers
Toward a Holistic Approach to Continual Model Merging
Hoang Phan, Sungmin Cha, Tung Lam Tran +1
We present a holistic framework for Continual Model Merging (CMM) that intervenes at three critical stages: pre-merging, during merging, and post-merging-to address two fundamental…
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
Shikai Qiu, Zixi Chen, Hoang Phan +2
Several recently introduced deep learning optimizers utilizing matrix-level preconditioning have shown promising speedups relative to the current dominant optimizer AdamW, particul…
Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
Hoang Phan, Victor Li, Qi Lei
Large language models (LLMs) have revolutionized natural language processing with their ability to generate coherent and contextually relevant text. However, their deployment raise…
Performative Risk Control: Calibrating Models for Reliable Deployment under Performativity
Victor Li, Baiting Chen, Yuzhen Mao +2
Calibrating blackbox machine learning models to achieve risk control is crucial to ensure reliable decision-making. A rich line of literature has been studying how to calibrate a m…