4 papers
Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
Shanbo Cheng, Yu Bao, Qian Cao +23
Multilingual translation stands as a challenging task for large language models (LLMs) to handle intricate language patterns and stilted translations that arise in automated transl…
DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Shuaijie She, Yu Bao, Yu Lu +7
We present DuPO, a dual learning-based preference optimization framework that generates annotation-free feedback via a generalized duality. DuPO addresses two key limitations: Rein…
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
Shimao Zhang, Zhejian Lai, Xiang Liu +5
Multilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual capabilities, which transfers the capabilities from the high-resource languages to…
R-PRM: Reasoning-Driven Process Reward Modeling
Shuaijie She, Junxiao Liu, Yifeng Liu +3
Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by eval…