6 papers
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
Zhiyi Mou, Jingyuan Yang, Zeheng Qian +6
While Large Language Models (LLMs) have powerful capabilities, they remain vulnerable to jailbreak attacks, which is a critical barrier to their safe web real-time application. Cur…
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
Ziyu Zhou, Yihang Wu, Jingyuan Yang +2
Black-Box prompt optimization methods have emerged as a promising strategy for refining input prompts to better align large language models (LLMs), thereby enhancing their task per…
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
Weixuan Wang, Jingyuan Yang, Wei Peng
Large language models (LLMs) have achieved remarkable performance across many tasks, yet aligning them with desired behaviors remains challenging. Activation intervention has emerg…
Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models
Jingyuan Yang, Bowen Yan, Rongjun Li +4
Unsafe prompts pose significant safety risks to large language models (LLMs). Existing methods for detecting unsafe prompts rely on data-driven fine-tuning to train guardrail model…
LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models
Jingyuan Yang, Rongjun Li, Weixuan Wang +3
Large Language Models (LLMs) often generate inconsistent responses when prompted with semantically equivalent paraphrased inputs. Recently, activation steering, a technique that mo…
Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach
Jingyuan Yang, Dapeng Chen, Yajing Sun +3
A Large Language Model (LLM) tends to generate inconsistent and sometimes contradictory outputs when presented with a prompt that has equivalent semantics but is expressed differen…