1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents
Chen Chen, Kim Young Il, Yuan Yang +7
Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
Chen Chen, Yuchen Sun, Jiaxin Gao +5
Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks. However, they remain vulnerable…
Neutralizing Backdoors through Information Conflicts for Large Language Models
Chen Chen, Yuchen Sun, Xueluan Gong +2
Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks, from understanding to reasoning.…
Hidden Data Privacy Breaches in Federated Learning
Xueluan Gong, Yuji Wang, Shuaike Li +5
Federated Learning (FL) emerged as a paradigm for conducting machine learning across broad and decentralized datasets, promising enhanced privacy by obviating the need for direct d…