large language models 1multi-turn jailbreak 1policy optimization 1reinforcement learning 1turn-level credit assignment 1
From the 1 of 9 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Junyoung Park, Sunghwan Park, Seongyong Ju +1
Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not how it unfolded. Two attacks t…
cs.AI2026
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
Junyoung Park, Dalton Jones, Matthew J Morse +3
We demonstrate that geometrically distinctive keys during LLM inference tend to have high attention scores. Based on the phenomenon we propose KeyDiff, a training-free KV cache evi…