activation analysis 1context calibration 1foundation models 1LLM agents 1policy size 1post-training compute allocation 1reinforcement learning 1reward feedback 1reward hacking 1safety monitoring 1search rollouts 1
From the 2 of 7 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
Patrick Wilhelm, Odej Kao
The paper investigates how to best allocate a fixed FLOP budget for reinforcement‑learning post‑training of foundation models, comparing larger policies, longer training, more sear…
cs.LG2026
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
Patrick Wilhelm, Odej Kao
In asynchronous federated learning (FL), client devices send updates to a central server at varying times based on their computational speed, often using stale versions of the glob…
cs.LG2026
Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding
Patrick Wilhelm, Inese Yilmaz, Odej Kao
Training large-scale Neural Networks requires substantial computational power and energy. Federated Learning enables distributed model training across geospatially distributed data…