2 papers
cs.LG2026
What do Reward Models Memorize?
Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallo…
cs.LG2026
Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning
Valentijn Oldenburg, Floris de Kam, Bente Zuijdam +4
Most parameter-efficient finetuning (PEFT) methods adapt weights or activations, thus leaving one of the key Transformer components unchanged: residual connections. This paper inve…