128 citations · 138 across the 4 of their papers we have counts for
3 papers · 1 filter
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu, Puxin Xu +12
Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and re…
Long-Term Planning and Situational Awareness in OpenAI Five
Jonathan Raiman, Susan Zhang, Filip Wolski
Understanding how knowledge about the world is represented within model-free deep reinforcement learning methods is a major challenge given the black box nature of its learning pro…
Neural Network Surgery with Sets
Jonathan Raiman, Susan Zhang, Christy Dennison
The cost to train machine learning models has been increasing exponentially, making exploration and research into the correct features and architecture a costly or intractable ende…