3 papers
cs.CL2025
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
Wen Yang, Junhong Wu, Chen Wang +2
Direct Preference Optimization (DPO) has become a prominent method for aligning Large Language Models (LLMs) with human preferences. While DPO has enabled significant progress in a…
cs.CL2024
Language Imbalance Driven Rewarding for Multilingual Self-improving
Wen Yang, Junhong Wu, Chen Wang +2
Large Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such…
cs.CL2024
Hit the Sweet Spot! Span-Level Ensemble for Large Language Models
Yangyifan Xu, Jianghao Chen, Junhong Wu +1
Ensembling various LLMs to unlock their complementary potential and leverage their individual strengths is highly valuable. Previous studies typically focus on two main paradigms:…