most citedCodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

416 citations · 488 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SE20224 cited

Execution-based Evaluation for Data Science Code Generation Models

Junjie Huang, Chenglong Wang, Jipeng Zhang +6

Code generation models can benefit data scientists' productivity by automatically generating code from context and text descriptions. An important measure of the modeling progress…

cs.SE202211 cited

Exploring and Evaluating Personalized Models for Code Generation

Andrei Zlotchevski, Dawn Drain, Alexey Svyatkovskiy +3

Large Transformer models achieved the state-of-the-art status for Natural Language Understanding tasks and are increasingly becoming the baseline model architecture for modeling so…

cs.SE2022

Generating Examples From CLI Usage: Can Transformers Help?

Roshanak Zilouchian Moghaddam, Spandan Garg, Colin B. Clement +2

Continuous evolution in modern software often causes documentation, tutorials, and examples to be out of sync with changing interfaces and frameworks. Relying on outdated documenta…

cs.SE202237 cited

Learning to Reduce False Positives in Analytic Bug Detectors

Anant Kharkar, Roshanak Zilouchian Moghaddam, Matthew Jin +4

Due to increasingly complex software design and rapid iterative development, code defects and security vulnerabilities are prevalent in modern software. In response, programmers re…

cs.LG202220 cited

Training and Evaluating a Jupyter Notebook Data Science Assistant

Shubham Chandel, Colin B. Clement, Guillermo Serrato +1

We study the feasibility of a Data Science assistant powered by a sequence-to-sequence transformer by training a new model JuPyT5 on all publicly available Jupyter Notebook GitHub…

cs.LG2021

Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy

Colin B. Clement, Shuai Lu, Xiaoyu Liu +5

Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks, setting high benchmarks for t…