A Convolutional Attention Network for Extreme Summarization of Source Code
arXiv:1602.03001
Abstract
Attention mechanisms in neural networks have proved useful for problems in which the input and output do not have fixed dimension. Often there exist features that are locally translation invariant and would be valuable for directing the model's attention, but previous attentional architectures are not constructed to learn such features specifically. We introduce an attentional neural network that employs convolution on the input tokens to detect local time-invariant and long-range topical attention features in a context-dependent way. We apply this architecture to the problem of extreme summarization of source code snippets into short, descriptive function name-like summaries. Using those features, the model sequentially generates a summary by marginalizing over two attention mechanisms: one that predicts the next summary token based on the attention weights of the input tokens and another that is able to copy a code token as-is directly into the summary. We demonstrate our convolutional attention neural network's performance on 10 popular Java projects showing that it achieves better performance compared to previous attentional mechanisms.
Code, data and visualization at http://groups.inf.ed.ac.uk/cup/codeattention/
References in corpus (4)
Cited by in corpus (28)
- Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review
- On the Feasibility of Transfer-learning Code Smells using Deep Learning
- Code Completion with Neural Attention and Pointer Networks
- CoaCor: Code Annotation for Code Retrieval with Reinforcement Learning
- StaQC: A Systematically Mined Question-Code Dataset from Stack Overflow
- Semantic Robustness of Models of Source Code
- Commit2Vec: Learning Distributed Representations of Code Changes
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Backdoors in Neural Models of Source Code
- Learning Continuous Semantic Representations of Symbolic Expressions
- On the Replicability and Reproducibility of Deep Learning in Software Engineering
- Incorporating Discrete Translation Lexicons into Neural Machine Translation
- Semantic Code Repair using Neuro-Symbolic Transformation Networks
- Toward Semi-Automatic Misconception Discovery Using Code Embeddings
- SmartPaste: Learning to Adapt Source Code
- Code Attention: Translating Code to Comments by Exploiting Domain Features
- A Survey on Machine Learning Techniques for Source Code Analysis
- Dialogue Act Recognition via CRF-Attentive Structured Network
- Reference-Aware Language Models
- Towards Neural Decompilation
- Deep Keyphrase Generation
- CORE: Automating Review Recommendation for Code Changes
- Fault Localization with Code Coverage Representation Learning
- Combining Code Embedding with Static Analysis for Function-Call Completion
- Automatically generating features for learning program analysis heuristics
- Pointing to Subwords for Generating Function Names in Source Code
- TAG : Type Auxiliary Guiding for Code Comment Generation
- Improving Automatic Source Code Summarization via Deep Reinforcement Learning