2 citations · 2 across the 3 of their papers we have counts for
3 papers
Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs
Kazuki Iwahana, Masaru Matsubayashi, Takuma Koyama +3
Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean inputs while producing attack…
Robust Backdoor Removal by Reconstructing Trigger-Activated Changes in Latent Representation
Kazuki Iwahana, Yusuke Yamasaki, Akira Ito +2
Backdoor attacks pose a critical threat to machine learning models, causing them to behave normally on clean data but misclassify poisoned data into a poisoned class. Existing defe…
First to Possess His Statistics: Data-Free Model Extraction Attack on Tabular Data
Masataka Tasumi, Kazuki Iwahana, Naoto Yanai +5
Model extraction attacks are a kind of attacks where an adversary obtains a machine learning model whose performance is comparable with one of the victim model through queries and…