activity
20242026
most citedThe Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

1 citations · 1 across the 12 of their papers we have counts for

collaborators

15 papers

cs.CL2026

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Chen Chen, Yaolin Chen, Xuehan Sun +5

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box a…

cs.CL2026

The Art of Mixology: Mixup-based Obfuscation for Privacy-Preserving Split Learning in Large Language Models

Chen Chen, Xiang Gao, Xianshun Wang +6

Split learning provides a practical paradigm for resource-constrained users to train Large Language Models (LLMs) by offloading computation-intensive layers to a server while keepi…

cs.RO2026

SoK: Security and Privacy of Foundation-Model-Powered Robots

Xueluan Gong, Chen Chen, Jinxin Liu +2

Foundation models are reshaping robotics by enabling robots to interpret open-ended instructions, reason over multimodal contexts, and operate in complex, open-world environments.…

cs.CR2026

Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations

Chen Chen, Yuchen Sun, Jiaxin Gao +4

Large language models (LLMs) are increasingly deployed in security-sensitive applications, yet remain vulnerable to backdoor attacks. However, existing backdoor defenses are diffic…

cs.CL20261 cited

The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

Chen Chen, Kim Young Il, Yuan Yang +7

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…

cs.CR2026

Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents

Kaiyu Zhou, Yongsen Zheng, Yicheng He +5

The agent--tool interaction loop is a critical attack surface for modern Large Language Model (LLM) agents. Existing denial-of-service (DoS) attacks typically function at the user-…