A Survey on Machine Learning Techniques for Source Code Analysis
arXiv:2110.09610
Abstract
The advancements in machine learning techniques have encouraged researchers to apply these techniques to a myriad of software engineering tasks that use source code analysis, such as testing and vulnerability detection. Such a large number of studies hinders the community from understanding the current research landscape. This paper aims to summarize the current knowledge in applied machine learning for source code analysis. We review studies belonging to twelve categories of software engineering tasks and corresponding machine learning techniques, tools, and datasets that have been applied to solve them. To do so, we conducted an extensive literature search and identified 479 primary studies published between 2011 and 2021. We summarize our observations and findings with the help of the identified studies. Our findings suggest that the use of machine learning techniques for source code analysis tasks is consistently increasing. We synthesize commonly used steps and the overall workflow for each task and summarize machine learning techniques employed. We identify a comprehensive list of available datasets and tools useable in this context. Finally, the paper discusses perceived challenges in this area, including the availability of standard datasets, reproducibility and replicability, and hardware resources.
References in corpus (12)
- Elixir: Effective object-oriented program repair
- Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges
- On the Efficiency of Test Suite based Program Repair: A Systematic Assessment of 16 Automated Repair Systems for Java Programs
- CoaCor: Code Annotation for Code Retrieval with Reinforcement Learning
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
- A parallel corpus of Python functions and documentation strings for automated code documentation and code generation
- Neural Program Repair by Jointly Learning to Localize and Repair
- A Survey of Deep Active Learning
- Latent Attention For If-Then Program Synthesis
- Towards Demystifying Dimensions of Source Code Embeddings
- DARVIZ: Deep Abstract Representation, Visualization, and Verification of Deep Learning Models
- An Experience Report on Machine Learning Reproducibility: Guidance for Practitioners and TensorFlow Model Garden Contributors