4 papers
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence
Yang Liu, Bin Chong, Wenkai Yang +9
Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to…
Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
Shan Yang, Yang Liu
Scaling cooperative multi-agent reinforcement learning (MARL) is fundamentally limited by cross-agent noise. When agents share a common reward, each agent's learning signal is comp…
From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions
Peiran Wang, Ying Li, Yuqiang Sun +3
The Model Context Protocol (MCP) has rapidly become a de facto standard for connecting LLM-based agents with external tools via reusable MCP servers. In practice, however, server s…
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
Ranjie Duan, Jiexi Liu, Xiaojun Jia +27
Large language models (LLMs) typically deploy safety mechanisms to prevent harmful content generation. Most current approaches focus narrowly on risks posed by malicious actors, of…