3 citations · 3 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
Zixuan Huang, Kecheng Huang, Lihao Yin +4
Large Language Models (LLMs) have gained significant traction in various applications, yet their capabilities present risks for both constructive and malicious exploitation. Despit…
cs.LG2025
Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor Attacks
Bowei He, Lihao Yin, Hui-Ling Zhen +4
The widespread deployment of pre-trained language models (PLMs) has exposed them to textual backdoor attacks, particularly those planted during the pre-training stage. These attack…