11 citations · 11 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
Shigeki Kusaka, Keita Saito, Mikoto Kudo +3
Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO a…
cs.CL2023★ 11 cited
Verbosity Bias in Preference Labeling by Large Language Models
Keita Saito, Akifumi Wachi, Koki Wataoka +1
In recent years, Large Language Models (LLMs) have witnessed a remarkable surge in prevalence, altering the landscape of natural language processing and machine learning. One key f…