low-resource languages 1meta-learning 1multilingual alignment 1preference learning 1reinforcement learning from human feedback 1
From the 1 of 6 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Meta-Learning Preferences for Multilingual LLM Alignment
Jiaying Lin, Seongho Son, Nam Phuong Tran +3
The paper introduces a meta-learning method that uses preference data from high-resource languages to quickly adapt large language models to low-resource languages with very few hu…
cs.CL2026
Overton Pluralistic Reinforcement Learning for Large Language Models
Yu Fu, Seongho Son, Ilija Bogunovic
Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism addresses this gap by generating responses with diverse perspecti…