works on

From the 1 of 26 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Meta-Learning Preferences for Multilingual LLM Alignment

Jiaying Lin, Seongho Son, Nam Phuong Tran +3

The paper introduces a meta-learning method that uses preference data from high-resource languages to quickly adapt large language models to low-resource languages with very few hu…

cs.CL2026

Overton Pluralistic Reinforcement Learning for Large Language Models

Yu Fu, Seongho Son, Ilija Bogunovic

Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism addresses this gap by generating responses with diverse perspecti…

cs.CL2026

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer +5

RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deployment requires reliable performance across…

cs.CL2025

This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs

Lorenz Wolf, Sangwoong Yoon, Ilija Bogunovic

Mixture of large language model (LLMs) Agents (MoA) architectures achieve state-of-the-art performance on prominent benchmarks like AlpacaEval 2.0 by leveraging the collaboration o…

cs.CL2024

Group Robust Preference Optimization in Reward-free RLHF

Shyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas +4

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data…