4 papers
RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing
Wenhui Zhang, Huiyu Xu, Zhibo Wang +4
Recent advancements in multi-model AI systems have leveraged LLM routers to reduce computational cost while maintaining response quality by assigning queries to the most appropriat…
Interpretable LLM Guardrails via Sparse Representation Steering
Zeqing He, Zhibo Wang, Huiyu Xu +3
Large language models (LLMs) exhibit impressive capabilities in generation tasks but are prone to producing harmful, misleading, or biased content, posing significant ethical and s…
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
Zeqing He, Zhibo Wang, Zhixuan Chu +4
Despite the outstanding performance of Large language Models (LLMs) in diverse tasks, they are vulnerable to jailbreak attacks, wherein adversarial prompts are crafted to bypass th…
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
Wenhui Zhang, Huiyu Xu, Zhibo Wang +3
Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees, and compara…