2 papers
cs.CR2026
CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs
Zhengxing Li, David J. Miller, Guangmingmei Yang +1
While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM i…
cs.LG2025
Inverting Trojans in LLMs
Zhengxing Li, Guangmingmei Yang, Jayaram Raghuram +2
While effective backdoor detection and inversion schemes have been developed for AIs used e.g. for images, there are challenges in "porting" these methods to LLMs. First, the LLM i…