From the 2 of 9 linked papers with an AI index.
9 papers
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
Patrick Wilhelm, Odej Kao
The paper investigates how to best allocate a fixed FLOP budget for reinforcement‑learning post‑training of foundation models, comparing larger policies, longer training, more sear…
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents
Patrick Wilhelm, Odej Kao
The paper investigates how internal activation signals, token entropy, and decision-context features can be used to monitor and mitigate reward‑hacking behavior in language‑model a…
Exploring Silent Data Corruption as a Reliability Challenge in LLM Training
Anton Altenbernd, Philipp Wiesner, Odej Kao
As Large Language Models (LLMs) scale in size and complexity, the consequences of failures during training become increasingly severe. A major challenge arises from Silent Data Cor…
Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation
Patrick Wilhelm, Odej Kao
In asynchronous federated learning (FL), client devices send updates to a central server at varying times based on their computational speed, often using stale versions of the glob…
Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding
Patrick Wilhelm, Inese Yilmaz, Odej Kao
Training large-scale Neural Networks requires substantial computational power and energy. Federated Learning enables distributed model training across geospatially distributed data…
Monitoring Emergent Reward Hacking During Generation via Internal Activations
Patrick Wilhelm, Thorsten Wittkopp, Odej Kao
Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has…