Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Spectral Scaling Laws of Muon
Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon. To…
cs.LG2026
Collaborative and Efficient Fine-tuning: Leveraging Task Similarity
Gagik Magakyan, Amirhossein Reisizadeh, Chanwoo Park +2
Adaptability has been regarded as a central feature in the foundation models, enabling them to effectively acclimate to unseen downstream tasks. Parameter-efficient fine-tuning met…
cs.LG2024
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
Kihyun Kim, Jiawei Zhang, Asuman Ozdaglar +1
Inverse Reinforcement Learning (IRL) and Reinforcement Learning from Human Feedback (RLHF) are pivotal methodologies in reward learning, which involve inferring and shaping the und…