3 papers
cs.CL2025
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
Terry Tong, Fei Wang, Zhe Zhao +1
This paper proposes a novel backdoor threat attacking the LLM-as-a-Judge evaluation regime, where the adversary controls both the candidate and evaluator model. The backdoored eval…
cs.LG2025
Unraveling Indirect In-Context Learning Using Influence Functions
Hadi Askari, Shivanshu Gupta, Terry Tong +3
In this work, we introduce a novel paradigm for generalized In-Context Learning (ICL), termed Indirect In-Context Learning. In Indirect ICL, we explore demonstration selection stra…
cs.CR2024
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
Qin Liu, Wenjie Mo, Terry Tong +4
The advancement of Large Language Models (LLMs) has significantly impacted various domains, including Web search, healthcare, and software development. However, as these models sca…