From the 1 of 9 linked papers with an AI index.
9 papers
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Zhengze Huang, Luyang Yu, Di Hong +5
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Weiwei Qi, Zefeng Wu, Zhilin Guo +5
Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained o…
DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
Zefeng Wu, Weiwei Qi, Jielong Chen +6
The paper introduces DataShield, a framework that detects risky fine‑tuning data for large language models by aligning safety‑critical semantic subspaces across multiple safety‑ali…
NonTextual Target Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng +6
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However,…
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
Churui Zeng, Weiwei Qi, Kedong Xiu +5
The rise of LLM agents introduces a new threat by enabling planning, coding, and even end-to-end execution of expert-level attack workflows. However, this threat remains underexplo…
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
Weiwei Qi, Zefeng Wu, Tianhang Zheng +4
Ensuring Large Language Model (LLM) safety is crucial, yet the lack of a clear understanding about safety mechanisms hinders the development of precise and reliable methodologies f…