activity
20242026
collaborators

8 papers

cs.MM2026

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

Yiming Sun, Chen Chen, Zifan Zhou +1

Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, they reach far more functionalities than CLI…

cs.AI2026

AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions

Chen Chen, Xueluan Gong, Ziyao Liu +3

AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems. With the rapid proliferation of AI and especially with the recent advancemen…

cs.CL2026

The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents

Chen Chen, Kim Young Il, Yuan Yang +7

Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…

cs.CL2025

Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution

Chen Chen, Yuchen Sun, Jiaxin Gao +5

Large language models (LLMs) have seen significant advancements, achieving superior performance in various Natural Language Processing (NLP) tasks. However, they remain vulnerable…

cs.CR2025

PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs

Xueluan Gong, Mingzhe Li, Yilin Zhang +5

Large Language Models (LLMs) have excelled in various tasks but are still vulnerable to jailbreaking attacks, where attackers create jailbreak prompts to mislead the model to produ…

cs.CR2025

Threats, Attacks, and Defenses in Machine Unlearning: A Survey

Ziyao Liu, Huanyi Ye, Chen Chen +2

Machine Unlearning (MU) has recently gained considerable attention due to its potential to achieve Safe AI by removing the influence of specific data from trained Machine Learning…