Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
Rui Wu, Yihao Quan, Zeru Shi +3
Safety-aligned Large Language Models (LLMs) still show two dominant failure modes: they are easily jailbroken, or they over-refuse harmless inputs that contain sensitive surface si…
cs.CL2024
Auto-Prompt Generation is Not Robust: Prompt Optimization Driven by Pseudo Gradient
Zeru Shi, Zhenting Wang, Yongye Su +5
While automatic prompt generation methods have recently received significant attention, their robustness remains poorly understood. In this paper, we introduce PertBench, a compreh…