collaborators

6 papers

cs.CR2026

SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails

Zhiyi Mou, Jingyuan Yang, Zeheng Qian +6

While Large Language Models (LLMs) have powerful capabilities, they remain vulnerable to jailbreak attacks, which is a critical barrier to their safe web real-time application. Cur…

cs.CL2025

Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow

Ziyu Zhou, Yihang Wu, Jingyuan Yang +2

Black-Box prompt optimization methods have emerged as a promising strategy for refining input prompts to better align large language models (LLMs), thereby enhancing their task per…

cs.CL2025

Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors

Weixuan Wang, Jingyuan Yang, Wei Peng

Large language models (LLMs) have achieved remarkable performance across many tasks, yet aligning them with desired behaviors remains challenging. Activation intervention has emerg…

cs.CL2025

Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models

Jingyuan Yang, Bowen Yan, Rongjun Li +4

Unsafe prompts pose significant safety risks to large language models (LLMs). Existing methods for detecting unsafe prompts rely on data-driven fine-tuning to train guardrail model…

cs.CL2025

LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models

Jingyuan Yang, Rongjun Li, Weixuan Wang +3

Large Language Models (LLMs) often generate inconsistent responses when prompted with semantically equivalent paraphrased inputs. Recently, activation steering, a technique that mo…

cs.CL2025

Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach

Jingyuan Yang, Dapeng Chen, Yajing Sun +3

A Large Language Model (LLM) tends to generate inconsistent and sometimes contradictory outputs when presented with a prompt that has equivalent semantics but is expressed differen…