papers

Publications (69)

cs.SE2023

Loghub: A Large Collection of System Log Datasets for AI-driven Log Analytics

Jieming Zhu, Shilin He, Pinjia He +2

Logs have been widely adopted in software system development and maintenance because of the rich runtime information they record. In recent years, the increase of software size and…

cs.CV2026

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation

Changyue Li, Jiaying Li, Youliang Yuan +3

Multimodal Large Language Models (MLLMs) are increasingly deployed in stateless systems, such as autonomous driving and robotics. This paper investigates a novel threat: Semantic-A…

cs.AI2026

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

Aoyang Fang, Yifan Yang, Jin'ao Shang +7

Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suff…

cs.AI2026

The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context

Xiaoyuan Liu, Tian Liang, Dongyang Ma +4

In the world of Harry Potter, when Dumbledore's mind is overburdened, he extracts memories into a Pensieve to be revisited later. In the world of AI, while we possess the Pensieve-…

cs.CL2025

Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Youliang Yuan, Wenxiang Jiao, Wenxuan Wang +5

This study addresses a critical gap in safety tuning practices for Large Language Models (LLMs) by identifying and tackling a refusal position bias within safety tuning data, which…

cs.SE2023

Incident-aware Duplicate Ticket Aggregation for Cloud Systems

Jinyang Liu, Shilin He, Zhuangbin Chen +10

In cloud systems, incidents are potential threats to customer satisfaction and business revenue. When customers are affected by incidents, they often request customer support servi…

cs.DB2019

Logzip: Extracting Hidden Structures via Iterative Clustering for Log Compression

Jinyang Liu, Jieming Zhu, Shilin He +3

System logs record detailed runtime information of software systems and are used as the main data source for many tasks around software engineering. As modern software systems are…

cs.SE2024

SPES: Towards Optimizing Performance-Resource Trade-Off for Serverless Functions

Cheryl Lee, Zhouruixing Zhu, Tianyi Yang +4

As an emerging cloud computing deployment paradigm, serverless computing is gaining traction due to its efficiency and ability to harness on-demand cloud resources. However, a sign…

cs.SE2025

A Goal-Driven Survey on Root Cause Analysis

Aoyang Fang, Haowen Yang, Haoze Dong +3

Root Cause Analysis (RCA) is a crucial aspect of incident management in large-scale cloud services. While the term root cause analysis or RCA has been widely used, different studie…

cs.SE2024

Prompting for Automatic Log Template Extraction

Junjielong Xu, Ruichun Yang, Yintong Huo +2

Log parsing, which involves log template extraction from semi-structured logs to produce structured logs, is the first and the most critical step in automated log analysis. However…

cs.CL2025

Towards Evaluating Proactive Risk Awareness of Multimodal Language Models

Youliang Yuan, Wenxiang Jiao, Yuejin Xie +5

Human safety awareness gaps often prevent the timely recognition of everyday risks. In solving this problem, a proactive safety artificial intelligence (AI) system would work bette…

cs.SE2026

Gleaner: A Semantically-Rich and Efficient Online Sampler for Microservice Diagnostics

Yifan Yang, Aoyang FANG, Songhan Zhang +1

Distributed tracing in microservices is critical for diagnostics but generates overwhelming data volumes, necessitating intelligent sampling. To maximize fidelity, state-of-the-art…

cs.SE2026

SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

Boxi Yu, Yang Cao, Yuzhong Zhang +9

The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals th…

cs.SE2023

Retromorphic Testing: A New Approach to the Test Oracle Problem

Boxi Yu, Qiuyang Mang, Qingshuo Guo +1

A test oracle serves as a criterion or mechanism to assess the correspondence between software output and the anticipated behavior for a given input set. In automated testing, blac…

cs.SE2022

Automated Testing of Image Captioning Systems

Boxi Yu, Zhiqing Zhong, Xinran Qin +3

Image captioning (IC) systems, which automatically generate a text description of the salient objects in an image (real or synthetic), have seen great progress over the past few ye…

cs.LG2025

CLEANet: Robust and Efficient Anomaly Detection in Contaminated Multivariate Time Series

Songhan Zhang, Yuanhao Lai, Pengfei Zheng +4

Multivariate time series (MTS) anomaly detection is essential for maintaining the reliability of industrial systems, yet real-world deployment is hindered by two critical challenge…

cs.SE2021

A Survey on Automated Log Analysis for Reliability Engineering

Shilin He, Pinjia He, Zhuangbin Chen +3

Logs are semi-structured text generated by logging statements in software source code. In recent decades, software logs have become imperative in the reliability assurance mechanis…

cs.SE2022

AEON: A Method for Automatic Evaluation of NLP Test Cases

Jen-tse Huang, Jianping Zhang, Wenxuan Wang +3

Due to the labor-intensive nature of manual test oracle construction, various automated testing techniques have been proposed to enhance the reliability of Natural Language Process…

cs.SE2021

Empirical Standards for Software Engineering Research

Paul Ralph, Nauman bin Ali, Sebastian Baltes +39

Empirical Standards are natural-language models of a scientific community's expectations for a specific kind of study (e.g. a questionnaire survey). The ACM SIGSOFT Paper and Peer…

cs.CL2023

BiasAsker: Measuring the Bias in Conversational AI System

Yuxuan Wan, Wenxuan Wang, Pinjia He +3

Powered by advanced Artificial Intelligence (AI) techniques, conversational AI systems, such as ChatGPT and digital assistants like Siri, have been widely deployed in daily life. H…

cs.SE2024

LILAC: Log Parsing using LLMs with Adaptive Parsing Cache

Zhihan Jiang, Jinyang Liu, Zhuangbin Chen +6

Log parsing transforms log messages into structured formats, serving as the prerequisite step for various log analysis tasks. Although a variety of log parsing approaches have been…

cs.SE2024

Unlocking the Power of Numbers: Log Compression via Numeric Token Parsing

Siyu Yu, Yifan Wu, Ying Li +1

Parser-based log compressors have been widely explored in recent years because the explosive growth of log volumes makes the compression performance of general-purpose compressors…

cs.SE2023

An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software

Wenxuan Wang, Jingyuan Huang, Jen-tse Huang +4

The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being…

cs.SE2023

Hue: A User-Adaptive Parser for Hybrid Logs

Junjielong Xu, Qiuai Fu, Zhouruixing Zhu +4

Log parsing, which extracts log templates from semi-structured logs and produces structured logs, is the first and the most critical step in automated log analysis. While existing…

cs.CL2025

Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs

Xiaoyuan Liu, Wenxuan Wang, Youliang Yuan +4

This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal co…

cs.SE2024

Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study

Yichen Li, Yintong Huo, Zhihan Jiang +5

Automated logging statement generation supports developers in documenting critical software runtime behavior. Given the great success in natural language generation and programming…

cs.CL2025

Scalable Supervising Software Agents with Patch Reasoner

Junjielong Xu, Boyin Tan, Xiaoyuan Liu +3

While large language model agents have advanced software engineering tasks, the unscalable nature of existing test-based supervision is limiting the potential improvement of data s…

cs.SE2019

Tools and Benchmarks for Automated Log Parsing

Jieming Zhu, Shilin He, Jinyang Liu +4

Logs are imperative in the development and maintenance process of many software systems. They record detailed runtime information that allows developers and support engineers to mo…

cs.SE2023

AutoLog: A Log Sequence Synthesis Framework for Anomaly Detection

Yintong Huo, Yichen Li, Yuxin Su +3

The rapid progress of modern computing systems has led to a growing interest in informative run-time logs. Various log-based anomaly detection techniques have been proposed to ensu…

cs.SE2026

LogPTR: Variable-Aware Log Parsing with Pointer Network

Yifan Wu, Bingxu Chai, Siyu Yu +4

Due to the sheer size of software logs, developers rely on automated log analysis. Log parsing, which parses semi-structured logs into a structured format, is a prerequisite of aut…

cs.SE2018

A Directed Acyclic Graph Approach to Online Log Parsing

Pinjia He, Jieming Zhu, Pengcheng Xu +2

Logs are widely used in modern software system management because they are often the only data accessible that record system events at runtime. In recent years, because of the ever…

cs.SE2024

Go Static: Contextualized Logging Statement Generation

Yichen Li, Yintong Huo, Renyi Zhong +6

Logging practices have been extensively investigated to assist developers in writing appropriate logging statements for documenting software behaviors. Although numerous automatic…

cs.CV2026

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

Jen-Tse Huang, Dasen Dai, Jen-Yuan Huang +7

Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, current Multimodal Large Language Models…

cs.SE2025

AL-Bench: A Benchmark for Automatic Logging

Boyin Tan, Junjielong Xu, Zhouruixing Zhu +1

Logging, the practice of inserting log statements into source code, is critical for improving software reliability. Recently, language model-based techniques have been developed to…

cs.SE2026

UniSage: A Unified and Post-Analysis-Aware Sampling for Microservices

Zhouruixing Zhu, Zhihan Jiang, Tianyi Yang +1

Traces and logs serve as the backbone of observability in microservice architectures, yet their sheer volume imposes prohibitive storage and computational burdens. To reduce overhe…

cs.SE2024

An Empirical Study on Package-Level Deprecation in Python Ecosystem

Zhiqing Zhong, Shilin He, Haoxuan Wang +3

Open-source software (OSS) plays a crucial role in modern software development. Utilizing OSS code can greatly accelerate software development, reduce redundancy, and enhance relia…

cs.SE2018

CARP: Context-Aware Reliability Prediction of Black-Box Web Services

Jieming Zhu, Pinjia He, Qi Xie +2

Reliability prediction is an important task in software reliability engineering, which has been widely studied in the last decades. However, modelling and predicting user-perceived…

cs.SE2025

Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware Benchmark

Aoyang Fang, Songhan Zhang, Yifan Yang +7

While cloud-native microservice architectures have revolutionized software development, their inherent operational complexity makes failure Root Cause Analysis (RCA) a critical yet…

cs.SE2024

On the Influence of Data Resampling for Deep Learning-Based Log Anomaly Detection: Insights and Recommendations

Xiaoxue Ma, Huiqi Zou, Pinjia He +4

Numerous Deep Learning (DL)-based approaches have gained attention in software Log Anomaly Detection (LAD), yet class imbalance in training data remains a challenge, with anomalies…

cs.HC2025

Let AI Read First: Enhancing Reading Abilities for Individuals with Dyslexia through Artificial Intelligence

Sihang Zhao, Shoucong Carol Xiong, Bo Pang +2

Dyslexia, a neurological condition affecting approximately 12% of the global population, presents significant challenges to reading ability and quality of life. Existing assistive…

cs.CL2024

Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs

Sihang Zhao, Youliang Yuan, Xiaoying Tang +1

Multimodal Large Language Models (MLLMs) demonstrate a strong understanding of the real world and can even handle complex tasks. However, they still fail on some straightforward vi…

cs.CL2026

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

Sihang Zhao, Kangrui Yu, Youliang Yuan +2

Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where stu…

cs.SE2025

SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints

Zhiyu Fan, Kirill Vasilevski, Dayi Lin +6

The advancement of large language models (LLMs) and code agents has demonstrated significant potential to assist software engineering (SWE) tasks, such as autonomous issue resoluti…

cs.CL2024

GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Youliang Yuan, Wenxiang Jiao, Wenxuan Wang +4

Safety lies at the core of the development of Large Language Models (LLMs). There is ample work on aligning LLMs with human ethics and preferences, including data filtering in pret…

cs.SE2023

Validating Multimedia Content Moderation Software via Semantic Fusion

Wenxuan Wang, Jingyuan Huang, Chang Chen +5

The exponential growth of social media platforms, such as Facebook and TikTok, has revolutionized communication and content publication in human society. Users on these platforms c…

cs.RO2025

Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach

Piaopiao Jin, Qi Wang, Guokang Sun +3

Vision-language-action (VLA) models demonstrate strong generalization in robotic manipulation but face challenges in complex, real-world tasks. While supervised fine-tuning with de…

cs.SE2026

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Jialun Cao, Yuk-Kit Chan, Zixuan Ling +12

Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In…

cs.SE2025

UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench

Boxi Yu, Yuxuan Zhu, Pinjia He +1

The advent of Large Language Models (LLMs) has spurred the development of coding agents for real-world code generation. As a widely used benchmark for evaluating the code generatio…

cs.LG2026

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment

Changyue Li, Jiaming He, Youliang Yuan +4

Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, se…

cs.CL2026

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

Yusong Zhao, Yuejin Xie, Youliang Yuan +4

Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable multimodal large language mod…

cs.SE2025

Aligning the Objective of LLM-based Program Repair

Junjielong Xu, Ying Fu, Shin Hwei Tan +1

Large language models (LLMs) have achieved decent results on automated program repair (APR). However, the next token prediction training objective of decoder-only LLMs (e.g., GPT-4…

cs.SE2026

MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context Learning

Jianbo Yu, Yixuan Li, Hai Xu +5

Log parsing converts semi-structured logs into structured templates, forming a critical foundation for downstream analysis. Traditional syntax and semantic-based parsers often stru…

cs.SE2025

DynaCausal: Dynamic Causality-Aware Root Cause Analysis for Distributed Microservices

Songhan Zhang, Aoyang Fang, Yifan Yang +3

Cloud-native microservices enable rapid iteration and scalable deployment but also create complex, fast-evolving dependencies that challenge reliable diagnosis. Existing root cause…

cs.CL2023

Automated Testing and Improvement of Named Entity Recognition Systems

Boxi Yu, Yiyan Hu, Qiuyang Mang +2

Named entity recognition (NER) systems have seen rapid progress in recent years due to the development of deep neural networks. These systems are widely used in various natural lan…

cs.SE2020

Structure-Invariant Testing for Machine Translation

Pinjia He, Clara Meister, Zhendong Su

In recent years, machine translation software has increasingly been integrated into our daily lives. People routinely use machine translation for various applications, such as desc…

cs.CV2017

Semantically Consistent Image Completion with Fine-grained Details

Pengpeng Liu, Xiaojuan Qi, Pinjia He +3

Image completion has achieved significant progress due to advances in generative adversarial networks (GANs). Albeit natural-looking, the synthesized contents still lack details, e…

cs.CL2018

Testing Untestable Neural Machine Translation: An Industrial Case

Wujie Zheng, Wenyu Wang, Dian Liu +6

Neural Machine Translation (NMT) has been widely adopted recently due to its advantages compared with the traditional Statistical Machine Translation (SMT). However, an NMT system…

cs.SE2025

BackportBench: A Multilingual Benchmark for Automated Backporting of Patches

Zhiqing Zhong, Jiaming Huang, Pinjia He

Many modern software projects evolve rapidly to incorporate new features and security patches. It is important for users to update their dependencies to safer versions, but many st…

cs.AI2025

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards

Xiaoyuan Liu, Tian Liang, Zhiwei He +6

Large Language Models (LLMs) show great promise in complex reasoning, with Reinforcement Learning with Verifiable Rewards (RLVR) being a key enhancement strategy. However, a preval…

cs.SE2024

LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models

Yuxuan Wan, Wenxuan Wang, Yiliu Yang +5

We introduce LogicAsker, a novel approach for evaluating and enhancing the logical reasoning capabilities of large language models (LLMs) such as ChatGPT and GPT-4. Despite LLMs' p…

cs.SE2023

ROME: Testing Image Captioning Systems via Recursive Object Melting

Boxi Yu, Zhiqing Zhong, Jiaqi Li +3

Image captioning (IC) systems aim to generate a text description of the salient objects in an image. In recent years, IC systems have been increasingly integrated into our daily li…

cs.SE2026

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

Manyi Wang, Junjielong Xu, Pinjia He

The paper introduces PAIChecker, a multi‑agent system that automatically detects misalignments between pull requests and their linked issues in SWE‑bench‑style benchmarks, improvin…

#software engineering#large language models#benchmark validation#pr‑issue alignment
cs.SE2026

DeLog: An Efficient Log Compression Framework with Pattern Signature Synthesis

Siyu Yu, Yifan Wu, Junjielong Xu +8

Parser-based log compression, which separates static templates from dynamic variables, is a promising approach to exploit the unique structure of log data. However, its performance…

cs.CL2025

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao +5

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the s…

cs.SE2026

SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding

Boyin Tan, Haoning Deng, Junyuan Zhang +3

Large language model (LLM) research in software engineering has largely focused on tasks such as code generation and bug repair. In practice, teams often draft multiple candidate p…

cs.CL2026

Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards

Youliang Yuan, Qiuyang Mang, Jingbang Chen +7

In this paper, we observe that current models are susceptible to reward hacking, leading to a substantial overestimation of a model's reasoning ability. This is evidenced by a high…

cs.CL2023

MTTM: Metamorphic Testing for Textual Content Moderation Software

Wenxuan Wang, Jen-tse Huang, Weibin Wu +5

The exponential growth of social media platforms such as Twitter and Facebook has revolutionized textual communication and textual content publication in human society. However, th…

cs.CL2021

Testing Machine Translation via Referential Transparency

Pinjia He, Clara Meister, Zhendong Su

Machine translation software has seen rapid progress in recent years due to the advancement of deep neural networks. People routinely use machine translation software in their dail…

cs.CR2015

A Privacy-Preserving QoS Prediction Framework for Web Service Recommendation

Jieming Zhu, Pinjia He, Zibin Zheng +1

QoS-based Web service recommendation has recently gained much attention for providing a promising way to help users find high-quality services. To facilitate such recommendations,…