activity
20242026
most citedSafety at Scale: A Comprehensive Survey of Large Model and Agent Safety

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models

Jiale Ding, Xiang Zheng, Yutao Wu +5

As large language models (LLMs) are increasingly deployed as black-box components in real-world applications, red teaming has become essential for identifying potential risks. It t…

cs.LG2025

Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks

Yige Li, Jiabo He, Hanxun Huang +3

Backdoor attacks have become a significant threat to the pre-training and deployment of deep neural networks (DNNs). Although numerous methods for detecting and mitigating backdoor…

cs.LG2025

Optimizing Cross-Client Domain Coverage for Federated Instruction Tuning of Large Language Models

Zezhou Wang, Yaxin Du, Xingjun Ma +3

Federated domain-specific instruction tuning (FedDIT) for large language models (LLMs) aims to enhance performance in specialized domains using distributed private and limited data…

cs.LG2025

SIDE: Surrogate Conditional Data Extraction from Diffusion Models

Yunhao Chen, Shujie Wang, Difan Zou +1

As diffusion probabilistic models (DPMs) become central to Generative AI (GenAI), understanding their memorization behavior is essential for evaluating risks such as data leakage,…

cs.LG2025

CHASE: A Causal Hypergraph based Framework for Root Cause Analysis in Multimodal Microservice Systems

Ziming Zhao, Zhenwei Wang, Tiehua Zhang +7

In recent years, the widespread adoption of distributed microservice architectures within the industry has significantly increased the demand for enhanced system availability and r…

cs.LG2025

FedEGG: Federated Learning with Explicit Global Guidance

Kun Zhai, Yifeng Gao, Difan Zou +4

Federated Learning (FL) holds great potential for diverse applications owing to its privacy-preserving nature. However, its convergence is often challenged by non-IID data distribu…