3 papers
cs.CR2026
FreoStream:Enhancing Stream Guardrails via Future-Aware Reasoning and Safety-Aligned Optimization
Jianwei Wang, Guoyang Shen, Yanhong Wu +5
Stream guardrails enable token-level safety detection before full responses are generated. However, they often make overly conservative judgements and block those sensitive but saf…
cs.CR2026
Activation-Guided Local Editing for Jailbreaking Attacks
Jiecong Wang, Haoran Li, Hao Peng +4
Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbreak methods face significant drawbacks.…
cs.CR2025
GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
Haoran Li, Yulin Chen, Jingru Zeng +7
As large language models (LLMs) are increasingly integrated into numerous applications across various domains, LLMs' safety becomes a critical concern for both application develope…