1 paper · 1 filter
Hyomin Lee, Sangwoo Park, Yumin Choi +3
While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities tha…