Contrastive Code Representation Learning
arXiv:2007.04973 · doi:10.18653/v1/2021.emnlp-main.482
Abstract
Recent work learns contextual representations of source code by reconstructing tokens from their context. For downstream semantic understanding tasks like summarizing code in English, these representations should ideally capture program functionality. However, we show that the popular reconstruction-based BERT model is sensitive to source code edits, even when the edits preserve semantics. We propose ContraCode: a contrastive pre-training task that learns code functionality, not form. ContraCode pre-trains a neural network to identify functionally similar variants of a program among many non-equivalent distractors. We scalably generate these variants using an automated source-to-source compiler as a form of data augmentation. Contrastive pre-training improves JavaScript summarization and TypeScript type inference accuracy by 2% to 13%. We also propose a new zero-shot JavaScript code clone detection dataset, showing that ContraCode is both more robust and semantically meaningful. On it, we outperform RoBERTa by 39% AUROC in an adversarial setting and up to 5% on natural code.
In Proceedings of EMNLP 2021. 19 pages, 16 figures, 9 tables. Code available at https://github.com/parasj/contracode
References in corpus (15)
- Improved Baselines with Momentum Contrastive Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
- On Mutual Information in Contrastive Learning for Visual Representations
- LambdaNet: Probabilistic Type Inference using Graph Neural Networks
- SCELMo: Source Code Embeddings from Language Models
- Adversarial Robustness for Code
- Scaling Laws for Transfer
- COSET: A Benchmark for Evaluating Neural Program Embeddings
- Learning Blended, Precise Semantic Program Embeddings
- Synthetic Datasets for Neural Program Synthesis
- OptTyper: Probabilistic Type Inference by Optimising Logical and Natural Constraints
- Evaluation of Generalizability of Neural Program Analyzers under Semantic-Preserving Transformations
Cited by in corpus (9)
- Evaluating Large Language Models Trained on Code
- Contrastive Representation Learning: A Framework and Review
- CLEAR: Contrastive Learning for Sentence Representation
- Code Structure Guided Transformer for Source Code Summarization
- SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation
- On the Effectiveness of Transfer Learning for Code Search
- DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
- ProtoTransformer: A Meta-Learning Approach to Providing Student Feedback
- How could Neural Networks understand Programs?