most citedFace It Yourselves: An LLM-Based Two-Stage Strategy to Localize Configuration Errors via Logs

25 citations · 27 across the 7 of their papers we have counts for

collaborators

7 papers

cs.SE2025

Generating High-Quality Datasets for Code Editing via Open-Source Language Models

Zekai Zhang, Mingwei Liu, Zhenxi Chen +7

Code editing plays a vital role in software engineering, requiring developers to adjust existing code according to natural language instructions while keeping functionality intact…

cs.SE2025

ConfLogger: Enhance Systems' Configuration Diagnosability through Configuration Logging

Shiwen Shan, Yintong Huo, Yuxin Su +3

Modern configurable systems offer customization via intricate configuration spaces, yet such flexibility introduces pervasive configuration-related issues such as misconfigurations…

cs.SE2025

AnomalyGen: An Automated Semantic Log Sequence Generation Framework with LLM for Anomaly Detection

Xinyu Li, Yingtong Huo, Chenxi Mao +4

The scarcity of high-quality public log datasets has become a critical bottleneck in advancing log-based anomaly detection techniques. Current datasets exhibit three fundamental li…

cs.LG2025★ 1 cited

BACE-RUL: A Bi-directional Adversarial Network with Covariate Encoding for Machine Remaining Useful Life Prediction

Zekai Zhang, Dan Li, Shunyu Wu +4

Prognostic and Health Management (PHM) are crucial ways to avoid unnecessary maintenance for Cyber-Physical Systems (CPS) and improve system reliability. Predicting the Remaining U…

cs.SE2025

LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance

Jingwen Tan, Gopi Krishnan Rajbahadur, Zi Li +5

Dataset license compliance is a critical yet complex aspect of developing commercial AI products, particularly with the increasing use of publicly available datasets. Ambiguities i…

cs.LG2024★ 1 cited

GLA-DA: Global-Local Alignment Domain Adaptation for Multivariate Time Series

Gang Tu, Dan Li, Bingxin Lin +2

Unlike images and natural language tokens, time series data is highly semantically sparse, resulting in labor-intensive label annotations. Unsupervised and Semi-supervised Domain A…