3 citations · 3 across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
Shaotian Yan, Kaiyuan Liu, Chen Shen +6
In this report, we introduce DASD-4B-Thinking, a lightweight yet highly capable, fully open-source reasoning model. It achieves SOTA performance among open-source models of compara…
cs.LG2024
ROPO: Robust Preference Optimization for Large Language Models
Xize Liang, Chao Chen, Shuang Qiu +6
Preference alignment is pivotal for empowering large language models (LLMs) to generate helpful and harmless responses. However, the performance of preference alignment is highly s…