10 papers
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu +4
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…
NonTextual Target Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
Yan Hong, Wei Li, Kedong Xiu +6
Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online d…
LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models
Yan Hong, Kedong Xiu, Wei Li +6
We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
Churui Zeng, Weiwei Qi, Kedong Xiu +5
The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
Bo Lv, Zhiheng Xu, KeDong Xiu +4
As Mixture-of-Experts (MoE) architectures are increasingly adopted for scaling Large Language Models (LLMs), safety auditing becomes necessary to verify whether these models produc…