1 citations · 1 across the 8 of their papers we have counts for
Showing 2026 · cs.CRShow all
2 papers · 2 filters
cs.CR2026
FreoStream:Enhancing Stream Guardrails via Future-Aware Reasoning and Safety-Aligned Optimization
Jianwei Wang, Guoyang Shen, Yanhong Wu +5
Stream guardrails enable token-level safety detection before full responses are generated. However, they often make overly conservative judgements and block those sensitive but saf…
cs.CR2026
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
Junke Zhang, Jianwei Wang, Sishuo Chen +3
Jailbreak attacks on large language models (LLMs) aim to induce LLMs to produce content that they are expected to refuse. Automated black-box jailbreak generation is important for…