2 papers
cs.CR2026
Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization
Yu Cui, Ruiqing Yue, Tingyu Li +6
Safety defenses for large language models (LLMs) have been extensively studied, with existing approaches focusing on attack detection and refusal mechanisms. Such fixed-form direct…
cs.CR2026
Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing
Yu Cui, Ruiqing Yue, Hang Fu +6
With the wide adoption of personal AI assistants such as OpenClaw, privacy leakage in user interaction contexts with large language model (LLM) agents has become a critical issue.…