Semantic-Enhanced Indirect Call Analysis with Large Language Models
arXiv:2408.04344 · doi:10.1145/3691620.3695016
Abstract
In contemporary software development, the widespread use of indirect calls to achieve dynamic features poses challenges in constructing precise control flow graphs (CFGs), which further impacts the performance of downstream static analysis tasks. To tackle this issue, various types of indirect call analyzers have been proposed. However, they do not fully leverage the semantic information of the program, limiting their effectiveness in real-world scenarios. To address these issues, this paper proposes Semantic-Enhanced Analysis (SEA), a new approach to enhance the effectiveness of indirect call analysis. Our fundamental insight is that for common programming practices, indirect calls often exhibit semantic similarity with their invoked targets. This semantic alignment serves as a supportive mechanism for static analysis techniques in filtering out false targets. Notably, contemporary large language models (LLMs) are trained on extensive code corpora, encompassing tasks such as code summarization, making them well-suited for semantic analysis. Specifically, SEA leverages LLMs to generate natural language summaries of both indirect calls and target functions from multiple perspectives. Through further analysis of these summaries, SEA can determine their suitability as caller-callee pairs. Experimental results demonstrate that SEA can significantly enhance existing static analysis methods by producing more precise target sets for indirect calls.
Accepted by ASE'24
References in corpus (15)
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- GPTScan: Detecting Logic Vulnerabilities in Smart Contracts by Combining GPT with Program Analysis
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Prompt Injection attack against LLM-integrated Applications
- Yi: Open Foundation Models by 01.AI
- AutoPruner: Transformer-Based Call Graph Pruning
- CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context
- Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
- Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- LLMDFA: Analyzing Dataflow in Code with Large Language Models
- Together We Go Further: LLMs and IDE Static Analysis for Extract Method Refactoring
- Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?