5 papers
Improved Bounds for Private and Robust Alignment
Wenqian Weng, Yi He, Xingyu Zhou
In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on the suboptimality gap in both offline and…
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
Zhang Wei, Hanxuan Chen, Peilu Hu +19
Red-teaming is becoming a central part of large language model (LLM) safety evaluation, yet current practice still relies heavily on expert-written prompts or fixed benchmark suite…
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
Zishan Bai, Hanxuan Chen, Jiayi Gu +9
Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and metrics arrive faster than operators can…
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Ziqian Bi, Yinzhi Wang, Tianyang Wang +6
Long Chain-of-Thought (CoT) traces can improve reasoning accuracy, but repeatedly generating them is costly for smaller or latency-constrained language models. This paper studies a…
SquarePO: Differentially Private and Robust -Preference Optimization in Offline Direct Alignment
Xingyu Zhou, Yulian Wu, Wenqian Weng +1
In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To th…