CodeBERT: A Pre-Trained Model for Programming and Natural Languages
arXiv:2002.08155
Abstract
We present CodeBERT, a bimodal pre-trained model for programming language (PL) and nat-ural language (NL). CodeBERT learns general-purpose representations that support downstream NL-PL applications such as natural language codesearch, code documentation generation, etc. We develop CodeBERT with Transformer-based neural architecture, and train it with a hybrid objective function that incorporates the pre-training task of replaced token detection, which is to detect plausible alternatives sampled from generators. This enables us to utilize both bimodal data of NL-PL pairs and unimodal data, where the former provides input tokens for model training while the latter helps to learn better generators. We evaluate CodeBERT on two NL-PL applications by fine-tuning model parameters. Results show that CodeBERT achieves state-of-the-art performance on both natural language code search and code documentation generation tasks. Furthermore, to investigate what type of knowledge is learned in CodeBERT, we construct a dataset for NL-PL probing, and evaluate in a zero-shot setting where parameters of pre-trained models are fixed. Results show that CodeBERT performs better than previous pre-trained models on NL-PL probing.
Accepted to Findings of EMNLP 2020. 12 pages
References in corpus (4)
Cited by in corpus (59)
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Test Case Selection and Prioritization Using Machine Learning: A Systematic Literature Review
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Neural Transfer Learning for Repairing Security Vulnerabilities in C Code
- Contrastive Code Representation Learning
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
- Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models
- Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized Architectures
- ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing
- Unsupervised Translation of Programming Languages
- Exploring Software Naturalness through Neural Language Models
- MathBERT: A Pre-Trained Model for Mathematical Formula Understanding
- Many bioinformatics programming tasks can be automated with ChatGPT
- Graph-based, Self-Supervised Program Repair from Diagnostic Feedback
- On the Effectiveness of Transfer Learning for Code Search
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- NER-BERT: A Pre-trained Model for Low-Resource Entity Tagging
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLM
- The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
- Software Vulnerability Detection via Deep Learning over Disaggregated Code Graph Representation
- DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
- Neural Code Search Revisited: Enhancing Code Snippet Retrieval through Natural Language Intent
- On using distributed representations of source code for the detection of C security vulnerabilities
- Code Recommendation for Open Source Software Developers
- Learning to Find Usages of Library Functions in Optimized Binaries
- Leveraging Automated Unit Tests for Unsupervised Code Translation
- What Makes a Good TODO Comment?
- Literature review on vulnerability detection using NLP technology
- ProtoTransformer: A Meta-Learning Approach to Providing Student Feedback
- Multi-task Learning based Pre-trained Language Model for Code Completion
- Cascaded Fast and Slow Models for Efficient Semantic Code Search
- ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
- Code Summarization with Structure-induced Transformer
- Improving Code Summarization with Block-wise Abstract Syntax Tree Splitting
- PyTorrent: A Python Library Corpus for Large-scale Language Models
- Is a Single Model Enough? MuCoS: A Multi-Model Ensemble Learning for Semantic Code Search
- Learning C to x86 Translation: An Experiment in Neural Compilation
- Enriching Query Semantics for Code Search with Reinforcement Learning
- GraphCode2Vec: Generic Code Embedding via Lexical and Program Dependence Analyses
- A Comparison of Code Embeddings and Beyond
- Cocktail: Leveraging Ensemble Learning for Optimized Model Serving in Public Cloud
- How could Neural Networks understand Programs?
- SYNFIX: Automatically Fixing Syntax Errors using Compiler Diagnostics
- AugmentedCode: Examining the Effects of Natural Language Resources in Code Retrieval Models
- NaturalCC: A Toolkit to Naturalize the Source Code Corpus
- Exploiting Token and Path-based Representations of Code for Identifying Security-Relevant Commits
- Distilling Transformers for Neural Cross-Domain Search
- Linguacodus: A Synergistic Framework for Transformative Code Generation in Machine Learning Pipelines
- Using Document Similarity Methods to create Parallel Datasets for Code Translation
- Searching for Replacement Classes
- CodeQA: A Question Answering Dataset for Source Code Comprehension
- Universal Representation for Code
- FACOS: Finding API Relevant Contents on Stack Overflow with Semantic and Syntactic Analysis
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy