3 papers
cs.CR2026
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
Junyoung Park, Seongyong Ju, Sunghwan Park +1
As Large Language Models evolve for user convenience, vulnerability to jailbreak attacks continues to be reported despite ongoing efforts in safety training. Traditional jailbreak…
cs.AI2026
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Junyoung Park, Sunghwan Park, Seongyong Ju +1
Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not how it unfolded. Two attacks t…
cs.CR2026
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
Taein Lim, Seongyong Ju, Munhyeok Kim +2
Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit…