papers

Publications (312)

cs.CL2019

Question Relatedness on Stack Overflow: The Task, Dataset, and Corpus-inspired Models

Amirreza Shirani, Bowen Xu, David Lo +2

cs.SE2026

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning

Guoqing Wang, Chengran Yang, Xiaoxuan Zhou +4

cs.SE2021

Accessibility in Software Practice: A Practitioner's Perspective

Tingting Bi, Xin Xia, David Lo +3

cs.SE2024

Large Language Model for Vulnerability Detection and Repair: Literature Review and the Road Ahead

Xin Zhou, Sicong Cao, Xiaobing Sun +1

cs.SE2024

DAppSCAN: Building Large-Scale Datasets for Smart Contract Weaknesses in DApp Projects

Zibin Zheng, Jianzhong Su, Jiachi Chen +3

cs.SE2021

CrossASR++: A Modular Differential Testing Framework for Automatic Speech Recognition

Muhammad Hilmi Asyrofi, Zhou Yang, David Lo

cs.SE2024

Assessing AI Detectors in Identifying AI-Generated Code: Implications for Education

Wei Hung Pan, Ming Jie Chok, Jonathan Leong Shan Wong +6

cs.CR2024

Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection Systems

Sicong Cao, Xiaobing Sun, Xiaoxue Wu +4

cs.SE2020

Automating App Review Response Generation

Cuiyun Gao, Jichuan Zeng, Xin Xia +3

cs.SE2025

Bamboo: LLM-Driven Discovery of API-Permission Mappings in the Android Framework

Han Hu, Wei Minn, Yonghui Liu +6

cs.LG2019

TreeCaps: Tree-Structured Capsule Networks for Program Source Code Processing

Vinoj Jayasundara, Nghi Duy Quoc Bui, Lingxiao Jiang +1

cs.SE2020

Scalable Online Vetting of Android Apps for Measuring Declared SDK Versions and Their Consistency with API Calls

Daoyuan Wu, Debin Gao, David Lo

cs.SE2025

Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks

Xing Hu, Feifei Niu, Junkai Chen +5

cs.SE2024

VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection

Xin-Cheng Wen, Xinchen Wang, Yujia Chen +3

cs.SE2026

Semantic Drift in Bug Resolution: How Behavioral Signals Propagate from Reports to Tests and Patches

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang +7

cs.SE2020

Code2Que: A Tool for Improving Question Titles from Mined Code Snippets in Stack Overflow

Zhipeng Gao, Xin Xia, David Lo +2

cs.SE2019

Automatic Generation of Pull Request Descriptions

Zhongxin Liu, Xin Xia, Christoph Treude +2

cs.SE2024

Towards Better Comprehension of Breaking Changes in the NPM Ecosystem

Dezhen Kong, Jiakun Liu, Lingfeng Bao +1

cs.SE2025

Less is More: On the Importance of Data Quality for Unit Test Generation

Junwei Zhang, Xing Hu, Shan Gao +3

cs.CR2026

Virtualization-based Penetration Testing Study for Detecting Accessibility Abuse Vulnerabilities in Banking Apps in East and Southeast Asia

Wei Minn, Phong Phan, Vikas K. Malviya +6

cs.CR2026

Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?

Yikun Li, Ngoc Tan Bui, Ting Zhang +16

cs.MA2026

Agent System Operations: Categorization, Challenges, and Future Directions

Zexin Wang, Changhua Pei, Yuanhao Liu +10

cs.SE2025

Fixseeker: An Empirical Driven Graph-based Approach for Detecting Silent Vulnerability Fixes in Open Source Software

Yiran Cheng, Ting Zhang, Lwin Khin Shar +6

cs.SE2022

CodeMatcher: Searching Code Based on Sequential Semantics of Important Query Words

Chao Liu, Xin Xia, David Lo +3

cs.SE2022

Revisiting Neuron Coverage Metrics and Quality of Deep Neural Networks

Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi +1

cs.SE2020

Opportunities and Challenges in Code Search Tools

Chao Liu, Xin Xia, David Lo +3

cs.LG2024

BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets

Chen Gong, Zhou Yang, Yunpeng Bai +8

cs.CL2025

TACLR: A Scalable and Efficient Retrieval-based Method for Industrial Product Attribute Value Identification

Yindu Su, Huike Zou, Lin Sun +7

cs.SI2017

GitHub and Stack Overflow: Analyzing Developer Interests Across Multiple Social Collaborative Platforms

Roy Ka-Wei Lee, David Lo

cs.SE2020

Characterization and Automatic Update of Deprecated Machine-Learning API Usages

Stefanus Agus Haryono, Ferdian Thung, David Lo +2

cs.SE2025

How do Machine Learning Models Change?

Joel Castaño, Rafael Cabañas, Antonio Salmerón +2

cs.SE2026

An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning

Yikun Li, Jinfeng Jiang, Ting Zhang +7

cs.LG2023

A Study of Variable-Role-based Feature Enrichment in Neural Models of Code

Aftab Hussain, Md Rafiqul Islam Rabin, Bowen Xu +2

cs.SE2019

PatchNet: A Tool for Deep Patch Classification

Thong Hoang, Julia Lawall, Richard J. Oentaryo +2

cs.SE2025

Let the Trial Begin: A Mock-Court Approach to Vulnerability Detection using LLM-Based Agents

Ratnadira Widyasari, Martin Weyssow, Ivana Clairine Irsan +6

cs.SI2018

Wisdom in Sum of Parts: Multi-Platform Activity Prediction in Social Collaborative Sites

Roy Ka-Wei Lee, David Lo

cs.SE2022

Duplicate Bug Report Detection: How Far Are We?

Ting Zhang, DongGyun Han, Venkatesh Vinayakarao +5

cs.SE2023

Invalidator: Automated Patch Correctness Assessment via Semantic and Syntactic Reasoning

Thanh Le-Cong, Duc-Minh Luong, Xuan Bach D. Le +4

cs.SE2023

Software Architecture in Practice: Challenges and Opportunities

Zhiyuan Wan, Yun Zhang, Xin Xia +2

cs.SE2021

BiasFinder: Metamorphic Test Generation to Uncover Bias for Sentiment Analysis Systems

Muhammad Hilmi Asyrofi, Zhou Yang, Imam Nur Bani Yusuf +3

cs.SE2024

Large Language Models for Software Engineering: A Systematic Literature Review

Xinyi Hou, Yanjie Zhao, Yue Liu +7

cs.SE2024

PPT4J: Patch Presence Test for Java Binaries

Zhiyuan Pan, Xing Hu, Xin Xia +3

cs.SE2025

On the Diffusion of Test Smells in LLM-Generated Unit Tests

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang +5

cs.SE2026

Mut4All: Fuzzing Compilers via LLM-Synthesized Mutators Learned from Bug Reports

Bo Wang, Pengyang Wang, Chong Chen +9

cs.SE2025

The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget

Dangfeng Pan, Zhensu Sun, Cenyuan Zhang +2

cs.SE2020

Smart Contract Repair

Xiao Liang Yu, Omar Al-Bataineh, David Lo +1

cs.CR2026

"Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors

Yue Liu, Yanjie Zhao, Yunbo Lyu +3

cs.SE2021

psc2code: Denoising Code Extraction from Programming Screencasts

Lingfeng Bao, Zhenchang Xing, Xin Xia +3

cs.SE2023

DexBERT: Effective, Task-Agnostic and Fine-grained Representation Learning of Android Bytecode

Tiezhu Sun, Kevin Allix, Kisub Kim +5

cs.SE2025

LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Xin Zhou, Martin Weyssow, Ratnadira Widyasari +7

cs.SE2020

Generating Question Titles for Stack Overflow from Mined Code Snippets

Zhipeng Gao, Xin Xia, John Grundy +2

cs.SE2026

Revisiting Vulnerability Patch Identification on Data in the Wild

Ivana Clairine Irsan, Ratnadira Widyasari, Ting Zhang +7

cs.SE2024

CUPID: Leveraging ChatGPT for More Accurate Duplicate Bug Report Detection

Ting Zhang, Ivana Clairine Irsan, Ferdian Thung +1

cs.SE2023

The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language Models

Xin Zhou, Kisub Kim, Bowen Xu +3

cs.SE2025

Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings

Alexandra González, Xavier Franch, David Lo +1

cs.SE2025

Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

Jianming Chang, Xin Zhou, Lulu Wang +2

cs.SE2021

DEFECTCHECKER: Automated Smart Contract Defect Detection by Analyzing EVM Bytecode

Jiachi Chen, Xin Xia, David Lo +3

cs.SE2020

AndroEvolve: Automated Android API Update with Data Flow Analysis and Variable Denormalization

Stefanus A. Haryono, Ferdian Thung, David Lo +5

cs.CR2024

{A New Hope}: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated Generation

Shidong Pan, Zhen Tao, Thong Hoang +7

cs.SE2026

How Agentic AI Coding Assistants Become the Attacker's Shell

Yue Liu, Yanjie Zhao, Yunbo Lyu +3

cs.SE2022

Natural Attack for Pre-trained Models of Code

Zhou Yang, Jieke Shi, Junda He +1

cs.SE2025

From Code to Courtroom: LLMs as the New Software Judges

Junda He, Jieke Shi, Terry Yue Zhuo +5

cs.SE2022

ITSS: Interactive Web-Based Authoring and Playback Integrated Environment for Programming Tutorials

Eng Lieh Ouh, Benjamin Kok Siew Gan, David Lo

cs.SE2026

Demystifying Solana Bots: From GitHub Blueprints to On-Chain Fingerprints

Xiaoye Zheng, Yujing Chen, Minghao Wu +5

The paper empirically investigates Solana blockchain bots by analyzing 586 GitHub repositories and 200 on-chain bot addresses, creating a taxonomy of bot types and describing their…

#solana#blockchain bots#on-chain analysis#cryptocurrency trading
cs.SE2026

Knowledge-Enhanced Agentic Vulnerability Repair

Sicong Cao, Hao Ma, Le Yu +8

cs.SE2025

BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Terry Yue Zhuo, Minh Chien Vu, Jenny Chim +30

cs.SE2026

Tail-aware N-version Machine Learning Models for Reliable API Recommendation

Aoi Matsuda, Fumio Machida, David Lo

cs.SE2022

Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?

Hong Jin Kang, Khai Loong Aw, David Lo

cs.SE2026

"Go Home Copilot, You're Drunk": Understanding Developer Responses to Agent-Generated Code Review Comments

Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy +2

cs.SE2024

Towards a Classification of Open-Source ML Models and Datasets for Software Engineering

Alexandra González, Xavier Franch, David Lo +1

cs.SE2023

Evaluating Pre-trained Language Models for Repairing API Misuses

Ting Zhang, Ivana Clairine Irsan, Ferdian Thung +3

cs.SE2022

Can Identifier Splitting Improve Open-Vocabulary Language Model of Code?

Jieke Shi, Zhou Yang, Junda He +2

cs.SE2016

Automated Inference of Software Library Usage Patterns

Mohamed Aymen Saied, Ali Ouni, Houari Sahraoui +3

cs.SE2022

SkipFuzz: Active Learning-based Input Selection for Fuzzing Deep Learning Libraries

Hong Jin Kang, Pattarakrit Rattanukul, Stefanus Agus Haryono +4

cs.CR2017

Mining Sandboxes for Linux Containers

Zhiyuan Wan, David Lo, Xin Xia +2

cs.LG2024

Regret-Based Defense in Adversarial Reinforcement Learning

Roman Belaire, Pradeep Varakantham, Thanh Nguyen +1

cs.SE2022

Static Inference Meets Deep Learning: A Hybrid Type Inference Approach for Python

Yun Peng, Cuiyun Gao, Zongjie Li +4

cs.SE2025

Zero-Shot Cross-Domain Code Search without Fine-Tuning

Keyu Liang, Zhongxin Liu, Chao Liu +3

cs.SE2025

Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness

Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang +5

cs.SE2025

Why Is My Transaction Risky? Understanding Smart Contract Semantics and Interactions in the NFT Ecosystem

Yujing Chen, Xuanming Liu, Zhiyuan Wan +4

cs.SE2024

Representation Learning for Stack Overflow Posts: How Far are We?

Junda He, Zhou Xin, Bowen Xu +6

cs.SE2022

Code Smells in Machine Learning Systems

Jiri Gesi, Siqi Liu, Jiawei Li +6

cs.SE2026

Multi-level Code Optimization via Mixture of Prompts

Yun Peng, Jun Wan, Jiakun Liu +3

cs.SE2017

A More Accurate Model for Finding Tutorial Segments Explaining APIs

He Jiang, Jingxuan Zhang, Xiaochen Li +2

eess.AS2023

ASDF: A Differential Testing Framework for Automatic Speech Recognition Systems

Daniel Hao Xian Yuen, Andrew Yong Chen Pang, Zhou Yang +3

cs.CR2022

VulCurator: A Vulnerability-Fixing Commit Detector

Truong Giang Nguyen, Thanh Le-Cong, Hong Jin Kang +2

cs.SE2026

To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair

Zhihao Lin, Junhua Zhu, Mingyi Zhou +5

cs.CR2023

A Closer Look at the Security Risks in the Rust Ecosystem

Xiaoye Zheng, Zhiyuan Wan, Yun Zhang +2

cs.CR2026

Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation

Yuchen Chen, Wei Cheng, Yuan Xiao +7

cs.SE2025

VERCATION: Precise Vulnerable Open-source Software Version Identification based on Static Analysis and LLM

Yiran Cheng, Ting Zhang, Lwin Khin Shar +6

cs.DC2021

Sage: Leveraging ML to Diagnose Unpredictable Performance in Cloud Microservices

Yu Gan, Mingyu Liang, Sundar Dev +2

cs.SE2022

PTM4Tag: Sharpening Tag Recommendation of Stack Overflow Posts with Pre-trained Models

Junda He, Bowen Xu, Zhou Yang +3

cs.SE2024

Bridging Expert Knowledge with Deep Learning Techniques for Just-In-Time Defect Prediction

Xin Zhou, DongGyun Han, David Lo

cs.SE2023

On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code

Martin Weyssow, Xin Zhou, Kisub Kim +2

cs.SE2022

AutoPRTitle: A Tool for Automatic Pull Request Title Generation

Ivana Clairine Irsan, Ting Zhang, Ferdian Thung +2

cs.SE2021

IncBL: Incremental Bug Localization

Zhou Yang, Jieke Shi, Shaowei Wang +1

cs.SE2026

What You Trust Is Insecure: Demystifying How Developers (Mis)Use Trusted Execution Environments in Practice

Yuqing Niu, Jieke Shi, Ruidong Han +4

cs.SE2022

Compressing Pre-trained Models of Code into 3 MB

Jieke Shi, Zhou Yang, Bowen Xu +2

cs.SE2024

Demystifying Faulty Code with LLM: Step-by-Step Reasoning for Explainable Fault Localization

Ratnadira Widyasari, Jia Wei Ang, Truong Giang Nguyen +2

cs.SE2026

Towards Fair Machine Learning Software: Understanding and Addressing Model Bias Through Counterfactual Thinking

Zichong Wang, Yang Zhou, David Lo +1