On the Effectiveness of Transfer Learning for Code Search
arXiv:2108.05890 · doi:10.1109/TSE.2022.3192755
Abstract
The Transformer architecture and transfer learning have marked a quantum leap in natural language processing, improving the state of the art across a range of text-based tasks. This paper examines how these advancements can be applied to and improve code search. To this end, we pre-train a BERT-based model on combinations of natural language and source code data and fine-tune it on pairs of StackOverflow question titles and code answers. Our results show that the pre-trained models consistently outperform the models that were not pre-trained. In cases where the model was pre-trained on natural language "and" source code data, it also outperforms an information retrieval baseline based on Lucene. Also, we demonstrated that the combined use of an information retrieval-based approach followed by a Transformer leads to the best results overall, especially when searching into a large search pool. Transfer learning is particularly effective when much pre-training data is available and fine-tuning data is limited. We demonstrate that natural language processing models based on the Transformer architecture can be directly applied to source code analysis tasks, such as code search. With the development of Transformer models designed more specifically for dealing with source code data, we believe the results of source code analysis tasks can be further improved.
Accepted for publication in the IEEE Transactions on Software Engineering (TSE)
References in corpus (17)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Cross-lingual Language Model Pretraining
- Evaluating Large Language Models Trained on Code
- Zero-Shot Learning Through Cross-Modal Transfer
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- CodeBERT: A Pre-Trained Model for Programming and Natural Languages
- Learning and Evaluating Contextual Embedding of Source Code
- Contrastive Code Representation Learning
- Knowledge Guided Text Retrieval and Reading for Open Domain Question Answering
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Multimodal Representation for Neural Code Search
- Gmail Smart Compose: Real-Time Assisted Writing
- Neural Code Search Evaluation Dataset
- Deep Transfer Learning for Source Code Modeling
- Learning Blended, Precise Semantic Program Embeddings
- Learning Scalable and Precise Representation of Program Semantics