Code Prediction by Feeding Trees to Transformers
arXiv:2003.13848
Abstract
We advance the state-of-the-art in the accuracy of code prediction (next token prediction) used in autocomplete systems. First, we report that using the recently proposed Transformer architecture even out-of-the-box outperforms previous neural and non-neural systems for code prediction. We then show that by making the Transformer architecture aware of the syntactic structure of code, we further increase the margin by which a Transformer-based system outperforms previous systems. With this, it outperforms the accuracy of an RNN-based system (similar to Hellendoorn et al. 2018) by 18.3%, the Deep3 system (Raychev et al 2016) by 14.1%, and an adaptation of Code2Seq (Alon et al., 2018) for code prediction by 14.4%. We present in the paper several ways of communicating the code structure to the Transformer, which is fundamentally built for processing sequence data. We provide a comprehensive experimental evaluation of our proposal, along with alternative design choices, on a standard Python dataset, as well as on a Facebook internal Python corpus. Our code and data preparation pipeline will be available in open source.
References in corpus (9)
- SmoothGrad: removing noise by adding noise
- Predicting Defective Lines Using a Model-Agnostic Technique
- Semantic Robustness of Models of Source Code
- Neural Program Repair by Jointly Learning to Localize and Repair
- Tree-Transformer: A Transformer-Based Method for Correction of Tree-Structured Data
- Structural Language Models of Code
- Tree-structured Attention with Hierarchical Accumulation
- COSET: A Benchmark for Evaluating Neural Program Embeddings
- Sequence Model Design for Code Completion in the Modern IDE
Cited by in corpus (12)
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Multimodal Representation for Neural Code Search
- Program Synthesis with Large Language Models
- Enquire One's Parent and Child Before Decision: Fully Exploit Hierarchical Structure for Self-Supervised Taxonomy Expansion
- DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
- Leveraging Automated Unit Tests for Unsupervised Code Translation
- Multi-task Learning based Pre-trained Language Model for Code Completion
- Learning Autocompletion from Real-World Datasets
- Combining Code Embedding with Static Analysis for Function-Call Completion
- How could Neural Networks understand Programs?
- NaturalCC: A Toolkit to Naturalize the Source Code Corpus
- Learning to Extend Program Graphs to Work-in-Progress Code