3 papers
cs.AI2025
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
Jirui Yang, Hengqi Guo, Zhihui Lu +6
Large language models often face a three-way trade-off among detection accuracy, inference latency, and deployment cost when used in real-world safety-sensitive applications. This…
cs.LG2025
N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
Zheyu Lin, Jirui Yang, Yukui Qiu +3
Evaluating the safety robustness of LLMs is critical for their deployment. However, mainstream Red Teaming methods rely on online generation and black-box output analysis. These ap…
cs.CR2025
CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation
Jirui Yang, Zheyu Lin, Zhihui Lu +6
Large language models (LLMs) are widely used for task understanding and action planning in embodied intelligence (EI) systems, but their adoption substantially increases vulnerabil…