most citedStable LM 2 1.6B Technical Report

8 citations · 14 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG20244 cited

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models

Alex Havrilla, Andrew Dai, Laura O'Mahony +17

Synthetic data generation with Large Language Models is a promising paradigm for augmenting natural data over a nearly infinite range of tasks. Given this variety, direct compariso…

cs.CL20241 cited

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

Zaid Alyafeai, Michael Pieler, Hannah Teufel +8

Large Language Models (LLMs) have shown impressive results in multiple domains of natural language processing (NLP) but are mainly focused on the English language. Recently, more L…

cs.CL20241 cited

Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training

Michael Pieler, Marco Bellagente, Hannah Teufel +9

Recently published work on rephrasing natural text data for pre-training LLMs has shown promising results when combining the original dataset with the synthetically rephrased data.…

cs.CL2024

Stable Code Technical Report

Nikhil Pinnaparaju, Reshinth Adithyan, Duy Phung +8

We introduce Stable Code, the first in our new-generation of code language models series, which serves as a general-purpose base code language model targeting code completion, reas…

cs.CL20248 cited

Stable LM 2 1.6B Technical Report

Marco Bellagente, Jonathan Tow, Dakota Mahan +16

We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to…