2 papers
cs.AI2026
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
Mohammad Beigi, Ming Jin, Junshan Zhang +2
Reinforcement Learning from Human Feedback (RLHF) remains vulnerable to reward hacking, where models exploit spurious correlations in learned reward models to achieve high scores w…
cs.LG2024
Communication-Efficient Federated Learning over Wireless Channels via Gradient Sketching
Vineet Sunil Gattani, Junshan Zhang, Gautam Dasarathy
Large-scale federated learning (FL) over wireless multiple access channels (MACs) has emerged as a crucial learning paradigm with a wide range of applications. However, its widespr…