activation analysis 1context calibration 1foundation models 1LLM agents 1policy size 1post-training compute allocation 1reinforcement learning 1reward feedback 1reward hacking 1safety monitoring 1search rollouts 1
From the 2 of 7 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference
Patrick Wilhelm, Thorsten Wittkopp, Odej Kao
Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks but come with substantial energy and computational costs, particularly in request-heavy scenar…
cs.CL2026
Monitoring Emergent Reward Hacking During Generation via Internal Activations
Patrick Wilhelm, Thorsten Wittkopp, Odej Kao
Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has…