4 papers
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
Kazutoshi Shinoda, Kosuke Nishida, Kyosuke Nishida
Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious features such as response leng…
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai +4
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models su…
Can LLMs Detect Their Own Hallucinations?
Sora Kadotani, Kosuke Nishida, Kyosuke Nishida
Large language models (LLMs) can generate fluent responses, but sometimes hallucinate facts. In this paper, we investigate whether LLMs can detect their own hallucinations. We form…
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
Kosuke Nishida, Kyosuke Nishida, Kuniko Saito
Loss spikes, a phenomenon in which the loss value diverges suddenly, is a fundamental issue in the pre-training of large language models. This paper supposes that the non-uniformit…