4 papers
Test-time reward-guided alignment of language models by importance sampling on pre-logit space
Sekitoshi Kanai, Tsukasa Yoshida, Hiroshi Takahashi +2
Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this paper, we propose a new test-time rewa…
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai +4
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models su…
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai +2
Contrastive language image pre-training (CLIP) is an essential component of building modern vision-language foundation models. While CLIP demonstrates remarkable zero-shot performa…
Transfer Learning with Pre-trained Conditional Generative Models
Shin'ya Yamaguchi, Sekitoshi Kanai, Atsutoshi Kumagai +2
Transfer learning is crucial in training deep neural networks on new target tasks. Current transfer learning methods always assume at least one of (i) source and target task label…