8 citations · 11 across the 3 of their papers we have counts for
3 papers
Stable Code Technical Report
Nikhil Pinnaparaju, Reshinth Adithyan, Duy Phung +8
We introduce Stable Code, the first in our new-generation of code language models series, which serves as a general-purpose base code language model targeting code completion, reas…
Teaching Large Language Models to Reason with Reinforcement Learning
Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6
Reinforcement Learning from Human Feedback (\textbf{RLHF}) has emerged as a dominant approach for aligning LLM outputs with human preferences. Inspired by the success of RLHF, we s…
Stable LM 2 1.6B Technical Report
Marco Bellagente, Jonathan Tow, Dakota Mahan +16
We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to…