11 papers
Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
Xujun Che, Yuchen Yuan, Weida Zhao +1
Error-penalized scoring rules ( for a correct answer, for a wrong one, for abstaining) are increasingly prescribed against hallucination: a rational agent facing such…
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization
Huilin Zhou, Jian Zhao, Yilu Zhong +7
Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static…
Visual Attention Reasoning via Hierarchical Search and Self-Verification
Wei Cai, Jian Zhao, Yuchen Yuan +4
Multimodal Large Language Models (MLLMs) frequently hallucinate due to their reliance on fragile, linear reasoning and weak visual grounding. We propose Visual Attention Reasoning…
Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
Yuxiang He, Jian Zhao, Yuchen Yuan +8
The exponential growth of digital content presents significant challenges for content safety. Current moderation systems, often based on single models or fixed pipelines, exhibit l…
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
Xiuyuan Chen, Jian Zhao, Yuxiang He +10
While the deployment of large language models (LLMs) in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based atta…
ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
Xin Zhang, Jiaming Chu, Jian Zhao +5
Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including au…