Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
UniMark: Artificial Intelligence Generated Content Identification Toolkit
Meilin Li, Ji He, Yi Yu +5
The rapid proliferation of Artificial Intelligence Generated Content has precipitated a crisis of trust and urgent regulatory demands. However, existing identification tools suffer…
cs.CR2025
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
Xiaoya Lu, Dongrui Liu, Yi Yu +2
Despite the rapid development of safety alignment techniques for LLMs, defending against multi-turn jailbreaks is still a challenging task. In this paper, we conduct a comprehensiv…