Code2Snapshot: Using Code Snapshots for Learning Representations of Source Code
arXiv:2111.01097 · doi:10.1109/ICMLA55696.2022.00140
Abstract
There are several approaches for encoding source code in the input vectors of neural models. These approaches attempt to include various syntactic and semantic features of input programs in their encoding. In this paper, we investigate Code2Snapshot, a novel representation of the source code that is based on the snapshots of input programs. We evaluate several variations of this representation and compare its performance with state-of-the-art representations that utilize the rich syntactic and semantic features of input programs. Our preliminary study on the utility of Code2Snapshot in the code summarization and code classification tasks suggests that simple snapshots of input programs have comparable performance to state-of-the-art representations. Interestingly, obscuring input programs have insignificant impacts on the Code2Snapshot performance, suggesting that, for some tasks, neural models may provide high performance by relying merely on the structure of input programs.
The 21st IEEE International Conference on Machine Learning and Applications (ICMLA'22)
References in corpus (8)
- On the Generalizability of Neural Program Models with respect to Semantic-Preserving Program Transformations
- A Literature Study of Embeddings on Source Code
- Embedding Java Classes with code2vec: Improvements from Variable Obfuscation
- Understanding Neural Code Intelligence Through Program Simplification
- Memorization and Generalization in Neural Code Intelligence Models
- Towards Demystifying Dimensions of Source Code Embeddings
- Syntax-Guided Program Reduction for Understanding Neural Code Intelligence Models
- Code2Image: Intelligent Code Analysis by Computer Vision Techniques and Application to Vulnerability Prediction