118 citations · 213 across the 3 of their papers we have counts for
3 papers
cs.LG2024
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Shang Liu, Yu Pan, Guanting Chen +1
Learning a reward model (RM) from human preferences has been an important component in aligning large language models (LLMs). The canonical setup of learning RMs from pairwise pref…
cs.SE2024★ 118 cited
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang +10
The rapid development of large language models has revolutionized code intelligence in software development. However, the predominance of closed-source models has restricted extens…
cs.CL2024★ 95 cited
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
DeepSeek-AI, :, Xiao Bi +85
The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions,…