SourcererCC: Scaling Code Clone Detection to Big Code
arXiv:1512.06448 · doi:10.1145/2884781.2884877
Abstract
Despite a decade of active research, there is a marked lack in clone detectors that scale to very large repositories of source code, in particular for detecting near-miss clones where significant editing activities may take place in the cloned code. We present SourcererCC, a token-based clone detector that targets three clone types, and exploits an index to achieve scalability to large inter-project repositories using a standard workstation. SourcererCC uses an optimized inverted-index to quickly query the potential clones of a given code block. Filtering heuristics based on token ordering are used to significantly reduce the size of the index, the number of code-block comparisons needed to detect the clones, as well as the number of required token-comparisons needed to judge a potential clone. We evaluate the scalability, execution time, recall and precision of SourcererCC, and compare it to four publicly available and state-of-the-art tools. To measure recall, we use two recent benchmarks, (1) a large benchmark of real clones, BigCloneBench, and (2) a Mutation/Injection-based framework of thousands of fine-grained artificial clones. We find SourcererCC has both high recall and precision, and is able to scale to a large inter-project repository (250MLOC) using a standard workstation.
Accepted for publication at ICSE'16 (preprint, unrevised)
References in corpus (1)
Cited by in corpus (28)
- VulDeePecker: A Deep Learning-Based System for Multiclass Vulnerability Detection
- Oreo: Detection of Clones in the Twilight Zone
- Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review
- Toxic Code Snippets on Stack Overflow
- CODIT: Code Editing with Tree-Based Neural Models
- Aroma: Code Recommendation via Structural Code Search
- Challenging Machine Learning-based Clone Detectors via Semantic-preserving Code Transformations
- Multimodal Representation for Neural Code Search
- Cross-Language Code Search using Static and Dynamic Analyses
- Active Learning of Discriminative Subgraph Patterns for API Misuse Detection
- Unprecedented Code Change Automation: The Fusion of LLMs and Transformation by Example
- SourcererCC and SourcererCC-I: Tools to Detect Clones in Batch mode and During Software Development
- Recommending Stack Overflow Posts for Fixing Runtime Exceptions using Failure Scenario Matching
- The Android Update Problem: An Empirical Study
- Using a Nearest-Neighbour, BERT-Based Approach for Scalable Clone Detection
- Jupyter Notebooks on GitHub: Characteristics and Code Clones
- Detecting Near Duplicates in Software Documentation
- Senatus -- A Fast and Accurate Code-to-Code Recommendation Engine
- MSCCD: Grammar Pluggable Clone Detection Based on ANTLR Parser Generation
- TransformCode: A Contrastive Learning Framework for Code Embedding via Subtree Transformation
- The Struggles of LLMs in Cross-lingual Code Clone Detection
- Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
- Using the Uniqueness of Global Identifiers to Determine the Provenance of Python Software Source Code
- Duplicated Code Pattern Mining in Visual Programming Languages
- Semantic Clone Detection via Probabilistic Software Modeling
- From Innovations to Prospects: What Is Hidden Behind Cryptocurrencies?
- Creative and Correct: Requesting Diverse Code Solutions from AI Foundation Models
- Unveiling Code Clone Patterns in Open Source VR Software: An Empirical Study