activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline

André V. Duarte, Xuying li, Bin Zeng +3

If we cannot inspect the training data of a large language model (LLM), how can we ever know what it has seen? We believe the most compelling evidence arises when the model itself…

cs.CL2025

The Resurgence of GCG Adversarial Attacks on Large Language Models

Yuting Tan, Xuying Li, Zhuo Li +2

Gradient-based adversarial prompting, such as the Greedy Coordinate Gradient (GCG) algorithm, has emerged as a powerful method for jailbreaking large language models (LLMs). In thi…

cs.CL2025

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

Huizhen Shu, Xuying Li, Qirui Wang +3

With the rapid proliferation of Natural Language Processing (NLP), especially Large Language Models (LLMs), generating adversarial examples to jailbreak LLMs remains a key challeng…

cs.CL2025

Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach

Xuying Li, Zhuo Li, Yuji Kosuga +1

Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conf…

cs.CL2025

Output Length Effect on DeepSeek-R1's Safety in Forced Thinking

Xuying Li, Zhuo Li, Yuji Kosuga +1

Large Language Models (LLMs) have demonstrated strong reasoning capabilities, but their safety under adversarial conditions remains a challenge. This study examines the impact of o…

cs.CL2024

Precision Knowledge Editing: Enhancing Safety in Large Language Models

Xuying Li, Zhuo Li, Yuji Kosuga +2

Large language models (LLMs) have demonstrated remarkable capabilities, but they also pose risks related to the generation of toxic or harmful content. This work introduces Precisi…