activity
20242026
most citedThe Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CR2026

Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations

Chen Chen, Yuchen Sun, Jiaxin Gao +4

Large language models (LLMs) are increasingly deployed in security-sensitive applications, yet remain vulnerable to backdoor attacks. However, existing backdoor defenses are diffic…

cs.CL20261 cited

The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

Chen Chen, Kim Young Il, Yuan Yang +7

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…

cs.AI2025

Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems

Jiaxin Gao, Chen Chen, Yanwen Jia +3

Large Language Models (LLMs) are increasingly being used to autonomously evaluate the quality of content in communication systems, e.g., to assess responses in telecom customer sup…

cs.CL2025

Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution

Chen Chen, Yuchen Sun, Jiaxin Gao +5

Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks. However, they remain vulnerable…

cs.CL2024

Neutralizing Backdoors through Information Conflicts for Large Language Models

Chen Chen, Yuchen Sun, Xueluan Gong +2

Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks, from understanding to reasoning.…

cs.CL2024

Hidden Data Privacy Breaches in Federated Learning

Xueluan Gong, Yuji Wang, Shuaike Li +5

Federated Learning (FL) emerged as a paradigm for conducting machine learning across broad and decentralized datasets, promising enhanced privacy by obviating the need for direct d…