2 papers
cs.LG2026
Test-time reward-guided alignment of language models by importance sampling on pre-logit space
Sekitoshi Kanai, Tsukasa Yoshida, Hiroshi Takahashi +2
Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this paper, we propose a new test-time rewa…
cs.LG2026
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai +4
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models su…