automated testing 1benchmarking 1large language model agents 1persistent memory 1personal agents 1red teaming 1safety 1safety auditing 1sycophancy 1vulnerability discovery 1
From the 2 of 14 linked papers with an AI index.
1 citations · 1 across the 11 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
Jiale Ding, Xiang Zheng, Yutao Wu +5
As large language models (LLMs) are increasingly deployed as black-box components in real-world applications, red teaming has become essential for identifying potential risks. It t…
cs.LG2025
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
Xiang Zheng, Xingjun Ma, Wei-Bin Lee +1
Red teaming has proven to be an effective method for identifying and mitigating vulnerabilities in Large Language Models (LLMs). Reinforcement Fine-Tuning (RFT) has emerged as a pr…