collaborators

6 papers

cs.CL2026

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

Xinjie Shen, Rongzhe Wei, Peizhi Niu +6

Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, increas…

cs.CR2026

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

Junran Wang, Xinjie Shen, Zehao Jin +1

As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments become…

cs.CL2026

Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention

Zehao Jin, Ruixuan Deng, Junran Wang +2

Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model p…

cs.CR2026

Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark

Xinjie Shen, Mufei Li, Pan Li

The deployment of Large Language Models (LLMs) in embodied agents creates an urgent need to measure their privacy awareness in the physical world. Existing evaluation methods, howe…

cs.HC2026

Behavioral Indicators of Overreliance During Interaction with Conversational Language Models

Chang Liu, Qinyi Zhou, Xinjie Shen +3

LLMs are now embedded in a wide range of everyday scenarios. However, their inherent hallucinations risk hiding misinformation in fluent responses, raising concerns about overrelia…

cs.CR2025

The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search

Rongzhe Wei, Peizhi Niu, Xinjie Shen +7

Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Existing approaches overwhelmingly operate within the p…