From the 1 of 5 linked papers with an AI index.
1 citations · 1 across the 5 of their papers we have counts for
5 papers
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Yongjian Guo, Wanlun Ma, Lingyu Shen +2
The paper introduces Routing-based On-Policy Distillation (ROPD), a method for safely realigning large language models that resists malicious prompt templates while preserving the…
Beyond Her: Safety Dynamics in Role-play AI Companions
Zehang Deng, Zhaoyang Xie, Changzhou Han +8
The film 'Her' pictured a future of love between humans and AI. That future has quietly emerged in the form of Role-play AI Companions (RACs), where emotionally responsive interact…
Five Queries Are Enough: Query-Efficient and Surrogate-Free Membership Inference Attacks on RAG via Entailment
Nguyen Linh Bao Nguyen, Wanlun Ma, Viet Vo +4
Retrieval-augmented generation (RAG) has become central to large language model (LLM) deployments, grounding responses in enterprise or proprietary data to reduce hallucinations. H…
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
Yongjian Guo, Puzhuo Liu, Wanlun Ma +5
The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, significantly enhancing their functionality.…
Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training
Jiaxuan Gao, Yongjian Guo, Zhong Guan +5
The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned World Models as generative simu…