collaborators

10 papers

cs.LG2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

Xinzhe Huang, Biwu Yao, Kedong Xiu +4

Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…

cs.CR2026

NonTextual Target Attack

Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6

Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…

cs.AI2026

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

Yan Hong, Wei Li, Kedong Xiu +6

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online d…

stat.ML2026

LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models

Yan Hong, Kedong Xiu, Wei Li +6

We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…

cs.CR2026

TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking

Churui Zeng, Weiwei Qi, Kedong Xiu +5

The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…

cs.CR2026

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

Bo Lv, Zhiheng Xu, KeDong Xiu +4

As Mixture-of-Experts (MoE) architectures are increasingly adopted for scaling Large Language Models (LLMs), safety auditing becomes necessary to verify whether these models produc…