Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming
Zelalem Abahana, David Evans, Satish Mahadevan Srinivasan +1
RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. We study this problem with a compact RLHF…
cs.LG2026
NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning
Michael Jerge, David Evans
This paper presents NoisyCoconut, a novel inference-time method that enhances large language model (LLM) reliability by manipulating internal representations. Unlike fine-tuning me…