1 citations · 1 across the 7 of their papers we have counts for
8 papers
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
Zhewen Tan, Wenhan Yu, Jianfeng Si +9
In recent years, safety risks associated with large language models have become increasingly prominent, highlighting the urgent need to mitigate the generation of toxic and harmful…
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
Wenrui Liu, Zixiang Liu, Elsie Dai +5
Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future tren…
ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents
Yilun Yao, Shan Huang, Elsie Dai +5
Large language models are increasingly deployed as research agents for deep search and long-horizon information seeking, yet their performance often degrades as interaction histori…
EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation
Zihang Li, Yuhang Wang, Yikun Zong +6
Chain-of-Thought (CoT) prompting has significantly enhanced the mathematical reasoning capabilities of Large Language Models. We find existing fine-tuning datasets frequently suffe…
"Even GPT Can Reject Me": Conceptualizing Abrupt Refusal Secondary Harm (ARSH) and Reimagining Psychological AI Safety with Compassionate Completion Standard (CCS)
Yang Ni, Tong Yang
Large Language Models (LLMs) and AI chatbots are increasingly used for emotional and mental health support due to their low cost, immediacy, and accessibility. However, when safety…
Beyond Invisibility: Learning Robust Visible Watermarks for Stronger Copyright Protection
Tianci Liu, Tong Yang, Quan Zhang +1
As AI advances, copyrighted content faces growing risk of unauthorized use, whether through model training or direct misuse. Building upon invisible adversarial perturbation, recen…