2 papers
cs.CL2026
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
Kazutoshi Shinoda, Kosuke Nishida, Kyosuke Nishida
Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious features such as response leng…
cs.LG2026
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai +4
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models su…