From the 1 of 6 linked papers with an AI index.
6 papers
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Weiwei Qi, Zefeng Wu, Zhilin Guo +5
Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…
DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
Zefeng Wu, Weiwei Qi, Jielong Chen +6
The paper introduces DataShield, a framework that detects risky fine‑tuning data for large language models by aligning safety‑critical semantic subspaces across multiple safety‑ali…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Kaicheng Shen, Lingyu Li, Wen Wu +3
AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that may accumulate over ti…
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
Liang Shan, Kaicheng Shen, Wen Wu +9
Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. T…
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
Churui Zeng, Weiwei Qi, Kedong Xiu +5
The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…
Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration
Junjie Xu, Xingjiao Wu, Zihao Zhang +6
AI agents that plan, retain memory across sessions, invoke external tools and act with partial autonomy are transforming human--AI collaboration. Research on affective computing, s…