From the 1 of 3 linked papers with an AI index.
1 citations · 1 across the 3 of their papers we have counts for
3 papers
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Yongjian Guo, Wanlun Ma, Lingyu Shen +2
The paper introduces Routing-based On-Policy Distillation (ROPD), a method for safely realigning large language models that resists malicious prompt templates while preserving the…
Beyond Her: Safety Dynamics in Role-play AI Companions
Zehang Deng, Zhaoyang Xie, Changzhou Han +8
The film 'Her' pictured a future of love between humans and AI. That future has quietly emerged in the form of Role-play AI Companions (RACs), where emotionally responsive interact…
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
Yongjian Guo, Puzhuo Liu, Wanlun Ma +5
The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, significantly enhancing their functionality.…