activity
20192022
most citedEvaluating Large Language Models Trained on Code

1.5k citations · 2.3k across the 5 of their papers we have counts for

collaborators

5 papers

cs.IR20229 cited

DIANES: A DEI Audit Toolkit for News Sources

Xiaoxiao Shang, Zhiyuan Peng, Qiming Yuan +4

Professional news media organizations have always touted the importance that they give to multiple perspectives. However, in practice the traditional approach to all-sides has favo…

cs.CL2022152 cited

Text and Code Embeddings by Contrastive Pre-Training

Arvind Neelakantan, Tao Xu, Raul Puri +22

Text embeddings are useful features in many applications such as semantic search and computing text similarity. Previous work typically trains models customized for different use c…

cs.LG20211.5k cited

Evaluating Large Language Models Trained on Code

Mark Chen, Jerry Tworek, Heewoo Jun +55

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex p…

cs.LG202121 cited

Asymmetric self-play for automatic goal discovery in robotic manipulation

OpenAI OpenAI, Matthias Plappert, Raul Sampedro +13

We train a single, goal-conditioned policy that can solve many robotic manipulation tasks, including tasks with previously unseen goals and objects. We rely on asymmetric self-play…

cs.LG2019635 cited

Solving Rubik's Cube with a Robot Hand

OpenAI, Ilge Akkaya, Marcin Andrychowicz +16

We demonstrate that models trained only in simulation can be used to solve a manipulation problem of unprecedented complexity on a real robot. This is made possible by two key comp…