Learning Natural Coding Conventions
arXiv:1402.4182 · doi:10.1145/2635868.2635883
Abstract
Every programmer has a characteristic style, ranging from preferences about identifier naming to preferences about object relationships and design patterns. Coding conventions define a consistent syntactic style, fostering readability and hence maintainability. When collaborating, programmers strive to obey a project's coding conventions. However, one third of reviews of changes contain feedback about coding conventions, indicating that programmers do not always follow them and that project members care deeply about adherence. Unfortunately, programmers are often unaware of coding conventions because inferring them requires a global view, one that aggregates the many local decisions programmers make and identifies emergent consensus on style. We present NATURALIZE, a framework that learns the style of a codebase, and suggests revisions to improve stylistic consistency. NATURALIZE builds on recent work in applying statistical natural language processing to source code. We apply NATURALIZE to suggest natural identifier names and formatting conventions. We present four tools focused on ensuring natural code during development and release management, including code review. NATURALIZE achieves 94% accuracy in its top suggestions for identifier names and can even transfer knowledge about conventions across projects, leveraging a corpus of 10,968 open source projects. We used NATURALIZE to generate 18 patches for 5 open source projects: 14 were accepted.
References in corpus (1)
Cited by in corpus (66)
- Mining Idioms from Source Code
- Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges
- Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code
- Predicting Defective Lines Using a Model-Agnostic Technique
- A Convolutional Attention Network for Extreme Summarization of Source Code
- Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review
- Structured Generative Models of Natural Source Code
- Parameter-Free Probabilistic API Mining across GitHub
- Typilus: Neural Type Hints
- IR2Vec: LLVM IR based Scalable Program Embeddings
- How Do I Refactor This? An Empirical Study on Refactoring Trends and Topics in Stack Overflow
- Automated Correction for Syntax Errors in Programming Assignments using Recurrent Neural Networks
- Learning Python Code Suggestion with a Sparse Pointer Network
- Language-Agnostic Representation Learning of Source Code from Structure and Context
- Unsupervised Translation of Programming Languages
- Code Vectors: Understanding Programs Through Embedded Abstracted Symbolic Traces
- Automatic Test Improvement with DSpot: a Study with Ten Mature Open-Source Projects
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Context2Name: A Deep Learning-Based Approach to Infer Natural Variable Names from Usage Contexts
- Does BLEU Score Work for Code Migration?
- Learning to Generate Corrective Patches using Neural Machine Translation
- Cross-Language Code Search using Static and Dynamic Analyses
- Program Synthesis with Large Language Models
- Modeling Vocabulary for Big Code Machine Learning
- ProGraML: Graph-based Deep Learning for Program Optimization and Analysis
- Learning API Usages from Bytecode: A Statistical Approach
- A large-scale comparative analysis of Coding Standard conformance in Open-Source Data Science projects
- Towards Demystifying Dimensions of Source Code Embeddings
- Learning to Represent Programs with Property Signatures
- DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
- Styler: learning formatting conventions to repair Checkstyle violations
- Recovering Variable Names for Minified Code with Usage Contexts
- DeepBugs: A Learning Approach to Name-based Bug Detection
- SmartPaste: Learning to Adapt Source Code
- End-to-End Prediction of Buffer Overruns from Raw Source Code via Neural Memory Networks
- What Do Developers Discuss about Code Comments?
- Syntax-Guided Program Reduction for Understanding Neural Code Intelligence Models
- Do Comments follow Commenting Conventions? A Case Study in Java and Python
- NeuDep: Neural Binary Memory Dependence Analysis
- Variable Name Recovery in Decompiled Binary Code using Constrained Masked Language Modeling
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
- Tailored Mutants Fit Bugs Better
- Code2Snapshot: Using Code Snapshots for Learning Representations of Source Code
- A Context-based Automated Approach for Method Name Consistency Checking and Suggestion
- Language Modelling for Source Code with Transformer-XL
- Towards a Model to Appraise and Suggest Identifier Names
- Deep Learning-Based Identification of Inconsistent Method Names: How Far Are We?
- Recommendation of Exception Handling Code in Mobile App Development
- Estimating defectiveness of source code: A predictive model using GitHub content
- On the "Naturalness" of Buggy Code
- DevReplay: Automatic Repair with Editable Fix Pattern
- Software Ethology: An Accurate, Resilient, and Cross-Architecture Binary Analysis Framework
- Recommending Variable Names for Extract Local Variable Refactorings
- Can We Benchmark Code Review Studies? A Systematic Mapping Study of Methodology, Dataset, and Metric
- Deep Generation of Coq Lemma Names Using Elaborated Terms
- Improving Semantic Consistency of Variable Names with Use-Flow Graph Analysis
- Deep Data Flow Analysis
- Do People Prefer "Natural" code?
- Roosterize: Suggesting Lemma Names for Coq Verification Projects Using Deep Learning
- Adoption and Evolution of Code Style and Best Programming Practices in Open-Source Projects
- Why Developers Refactor Source Code: A Mining-based Study
- Learning Program Component Order
- Using GGNN to recommend log statement level
- Abstraction Refinement Guided by a Learnt Probabilistic Model
- Are My Invariants Valid? A Learning Approach
- Video Game Development in a Rush: A Survey of the Global Game Jam Participants