collaborators

7 papers

cs.CL2026

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Mingyu Luo, Ming Deng, Zilang Qiu +8

Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read…

cs.CR2026

Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment

Yuchong Xie, Mingyu Luo, Zesen Liu +7

Coding agents powered by large language models are becoming central modules of modern IDEs, helping users perform complex tasks by invoking tools. While powerful, tool invocation o…

cs.CL2026

FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection

Y. H. Zhou, Z. M. Ma, Y. J. Zhou +12

SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user ac…

cs.CL2026

CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation

Sijin Sun, Liangbin Zhao, Jiaxiang Cai +3

Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including lexical transitions near utter…

cs.CR2026

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

Mingyu Luo, Zihan Zhang, Zesen Liu +7

LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers often trust evidence such as test result…

cs.CR2026

QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents

Yuchong Xie, Zesen Liu, Mingyu Luo +7

Modern coding agents integrated into IDEs orchestrate powerful tools and high-privilege system access, creating a high-stakes attack surface. Prior work on Indirect Prompt Injectio…