collaborators

11 papers

cs.CR2026

Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners

Xiaojun Jia, Jie Liao, Simeng Qin +5

Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely o…

cs.AI2026

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines

Yifei Ge, Weisong Sun, Jinkun Xiao +8

Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution environments, and the underlying system. For…

cs.CR2026

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

Ziwei Wang, Jing Chen, Ruichao Liang +6

Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive t…

cs.CR2026

Sealing the Audit-Runtime Gap for LLM Skills

Tingda Shen, Yebo Feng, Konglin Zhu +3

Large language model (LLM) ecosystems such as Claude Code and ChatGPT increasingly rely on skills: packages of natural-language instructions and executable tools. Once in the LLM's…

cs.CY2026

A Visionary Look at Vibe Researching

Yebo Feng, Yang Liu

Vibe researching is an emerging paradigm in which human researchers provide high-level direction and critical judgment while LLM-based agents handle the labor-intensive execution o…

cs.LG2026

Can Distillation Mitigate Backdoor Attacks in Pre-trained Encoders?

TIngxu Han, Wei Song, Weisong Sun +7

Self-Supervised Learning (SSL) has become a prominent paradigm for pre-training encoders to learning general-purpose representations from unlabeled data and releasing them on third…