5 papers
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
Yan Wang, Zhixuan Chu, Zihao Xue +9
Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful…
Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
Zihao Xue, Yan Wang, Zhen Bi +7
Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety control fundamentally differe…
Make LLM Learn to Synthesize from Streaming Experiences through Feedback
Zhenlin Hu, Yan Wang, Zhen Bi +7
Large language models (LLMs) have been widely adopted for synthetic data generation, significantly reducing annotation costs. However, most existing studies treat synthesis as a se…
Thought Purity: A Defense Framework For Chain-of-Thought Attack
Zihao Xue, Zhen Bi, Long Ma +7
Large Reasoning Models (LRMs) leverage Chain-of-Thought (CoT) reasoning to solve complex tasks, but this explicit reasoning process introduces a critical vulnerability: adversarial…
Your One-Stop Solution for AI-Generated Video Detection
Long Ma, Zihao Xue, Yan Wang +6
Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessit…