1 citations · 1 across the 16 of their papers we have counts for
11 papers · 1 filter
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
Yunhao Chen, Xin Wang, Yixu Wang +6
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior…
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
Jiaqi Yu, Xin Wang, Yixu Wang +4
Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate. Existing guardrails fall int…
Towards Context-Invariant Safety Alignment for Large Language Models
Yixu Wang, Yang Yao, Xin Wang +4
Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
Yunhao Chen, Xin Wang, Juncheng Li +5
Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimization primarily in the prompt space…
Mechanistic Origin of Moral Indifference in Language Models
Lingyu Li, Yan Teng, Yingchun Wang
Existing behavioral alignment techniques for Large Language Models (LLMs) often neglect the discrepancy between surface compliance and internal unaligned representations, leaving L…
LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
Zhiyuan Ning, Tianle Gu, Jiaxin Song +8
The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse rang…