most citedA.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code

1 citations · 1 across the 11 of their papers we have counts for

collaborators

12 papers

cs.SE2026

Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub

Yuli Cheng, Xiaoyu Zhang, Jiongchi Yu +3

GitHub plays a critical role in modern software supply chains, making its security an important research concern. Existing studies have primarily focused on CI/CD automation, colla…

cs.SE2026

Human in the Loop for Fuzz Testing: Literature Review and the Road Ahead

Jiongchi Yu, Xiaolin Wen, Sizhe Cheng +3

Fuzz testing is one of the most effective techniques for detecting bugs and vulnerabilities in software. However, as the basis of fuzz testing, automated heuristics often fail to u…

cs.AI2026

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

Yu Li, Qiang Hu, Yao Zhang +3

Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable collaborative reasoning capabilities but introduce new attack surfaces, such as the sleeper agent,…

cs.CL2026

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang +4

With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user…

cs.SE2025

Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs

Jian Wang, Xiaofei Xie, Qiang Hu +4

Automated Program Repair (APR) plays a critical role in enhancing the quality and reliability of software systems. While substantial progress has been made in Java-based APR, large…

cs.SE2025

AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis

Jiongchi Yu, Weipeng Jiang, Xiaoyu Zhang +3

Understanding software faults is essential for empirical research in software development and maintenance. However, traditional fault analysis, while valuable, typically involves m…