activity
20242026
most citedWhen Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR2026

From Compression to Accountability: Harmless Copyright Protection for Dataset Distillation

Yan Liang, Ziyuan Yang, Mengyu Sun +2

Large-scale datasets have been a key driving force behind the rapid progress of deep learning, but their storage, computational, and energy costs have become increasingly prohibiti…

cs.CR2026

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems

Yihao Zhang, Kai Wang, Jiangrong Wu +7

Large Language Models (LLMs) face prominent security risks from jailbreaking, a practice that manipulates models to bypass built-in security constraints and generate unethical or u…

cs.CR2026

AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems

Yihao Zhang, Zeming Wei, Xiaokun Luan +7

Autonomous LLM-based agents increasingly operate as long-running processes forming densely interconnected multi-agent ecosystems, whose security properties remain largely unexplore…

cs.CR2025

Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

Chengcan Wu, Zhixin Zhang, Mingqian Xu +2

Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent design in MAS introduces unique trustworthi…

cs.CR2025

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

Zeming Wei, Chengcan Wu, Meng Sun

Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In particular, they pose risks of generat…