4 papers
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
Zhang Wei, Hanxuan Chen, Peilu Hu +19
Red-teaming is becoming a central part of large language model (LLM) safety evaluation, yet current practice still relies heavily on expert-written prompts or fixed benchmark suite…
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
Zishan Bai, Hanxuan Chen, Jiayi Gu +9
Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and metrics arrive faster than operators can…
Improved Bounds for Private and Robust Alignment
Wenqian Weng, Yi He, Xingyu Zhou
In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on the suboptimality gap in both offline and…
SquarePO: Differentially Private and Robust -Preference Optimization in Offline Direct Alignment
Xingyu Zhou, Yulian Wu, Wenqian Weng +1
In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To th…