3 papers
cs.CL2025
Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models
Jingyuan Yang, Bowen Yan, Rongjun Li +4
Unsafe prompts pose significant safety risks to large language models (LLMs). Existing methods for detecting unsafe prompts rely on data-driven fine-tuning to train guardrail model…
cs.CL2025
LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models
Jingyuan Yang, Rongjun Li, Weixuan Wang +3
Large Language Models (LLMs) often generate inconsistent responses when prompted with semantically equivalent paraphrased inputs. Recently, activation steering, a technique that mo…
cs.CL2025
Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach
Jingyuan Yang, Dapeng Chen, Yajing Sun +3
A Large Language Model (LLM) tends to generate inconsistent and sometimes contradictory outputs when presented with a prompt that has equivalent semantics but is expressed differen…