6 papers
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
Wenlong Deng, Jiaji Huang, Kaan Ozkara +4
Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode through the geometry of reinfor…
MuonBP: Faster Muon via Block-Periodic Orthogonalization
Ahmed Khaled, Kaan Ozkara, Tao Yu +2
Gradient orthogonalization is a simple strategy that shows great utility in speeding up gradient descent. The Muon optimizer (Jordan, Jin, et al., 2024) combines gradient orthogona…
MEL: Multi-level Ensemble Learning for Resource-Constrained Environments
Krishna Praneet Gudipaty, Walid A. Hanafy, Kaan Ozkara +4
AI inference at the edge is becoming increasingly common for low-latency services. However, edge environments are power- and resource-constrained, and susceptible to failures. Conv…
SPIRE: Conditional Personalization for Federated Diffusion Generative Models
Kaan Ozkara, Ruida Zhou, Suhas Diggavi
Recent advances in diffusion models have revolutionized generative AI, but their sheer size makes on device personalization, and thus effective federated learning (FL), infeasible.…
Stochastic Rounding for LLM Training: Theory and Practice
Kaan Ozkara, Tao Yu, Youngsuk Park
As the parameters of Large Language Models (LLMs) have scaled to hundreds of billions, the demand for efficient training methods -- balancing faster computation and reduced memory…
ADEPT: Hierarchical Bayes Approach to Personalized Federated Unsupervised Learning
Kaan Ozkara, Bruce Huang, Ruida Zhou +1
Statistical heterogeneity of clients' local data is an important characteristic in federated learning, motivating personalized algorithms tailored to the local data statistics. Tho…