Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
Terry Tong, Fei Wang, Zhe Zhao +1
This paper proposes a novel backdoor threat attacking the LLM-as-a-Judge evaluation regime, where the adversary controls both the candidate and evaluator model. The backdoored eval…
cs.CL2024
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
Terry Tong, Jiashu Xu, Qin Liu +1
Large language models (LLMs) have acquired the ability to handle longer context lengths and understand nuances in text, expanding their dialogue capabilities beyond a single uttera…