1.6k citations · 1.9k across the 18 of their papers we have counts for
4 papers · 1 filter
TinyGSM: achieving >80% on GSM8k with small language models
Bingbin Liu, Sebastien Bubeck, Ronen Eldan +5
Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving…
Positional Description Matters for Transformers Arithmetic
Ruoqi Shen, Sébastien Bubeck, Ronen Eldan +3
Transformers, central to the successes in modern Natural Language Processing, often falter on arithmetic tasks despite their vast capabilities --which paradoxically include remarka…
Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16
We introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for…
Sparks of Artificial General Intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan +11
Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks,…