activity
20232026
most citedPAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs

3 citations · 5 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Chen Chen, Yaolin Chen, Xuehan Sun +5

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box a…

cs.CL2026

The Art of Mixology: Mixup-based Obfuscation for Privacy-Preserving Split Learning in Large Language Models

Chen Chen, Xiang Gao, Xianshun Wang +6

Split learning provides a practical paradigm for resource-constrained users to train Large Language Models (LLMs) by offloading computation-intensive layers to a server while keepi…

cs.CL2026★ 1 cited

The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

Chen Chen, Kim Young Il, Yuan Yang +7

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…

cs.CL2025

Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution

Chen Chen, Yuchen Sun, Jiaxin Gao +5

Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks. However, they remain vulnerable…

cs.CL2024

Neutralizing Backdoors through Information Conflicts for Large Language Models

Chen Chen, Yuchen Sun, Xueluan Gong +2

Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks, from understanding to reasoning.…

cs.CL2024

Hidden Data Privacy Breaches in Federated Learning

Xueluan Gong, Yuji Wang, Shuaike Li +5

Federated Learning (FL) emerged as a paradigm for conducting machine learning across broad and decentralized datasets, promising enhanced privacy by obviating the need for direct d…