5 papers
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu +4
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Zhengze Huang, Luyang Yu, Di Hong +5
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…
NonTextual Target Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…
Dynamic Jailbreaking Attack
Kedong Xiu, Yunhan Yang, Churui Zeng +6
Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a static optimization strategy. However, thi…
DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
Xinzhe Huang, Kedong Xiu, Tianhang Zheng +5
Recent research has focused on exploring the vulnerabilities of Large Language Models (LLMs), aiming to elicit harmful and/or sensitive content from LLMs. However, due to the insuf…