31 papers
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Yunhao Chen, Xin Wang, Yixu Wang +6
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior…
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Kaicheng Shen, Lingyu Li, Wen Wu +3
AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that may accumulate over ti…
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
Xinyang Liao, Lingyu Li, Huacan Liu +5
As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly important. Otherwise, such systems may rapidly…
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
Liang Shan, Kaicheng Shen, Wen Wu +9
Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. T…
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
Jiaqi Yu, Xin Wang, Yixu Wang +4
Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate. Existing guardrails fall int…
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
Leitao Yuan, Qinghua Mao, Daizong Liu +5
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate encoders can generalize to clo…