1 paper
Shigeki Kusaka, Keita Saito, Mikoto Kudo +3
Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO a…