2 papers
cs.CR2026
Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks
Yiyong Liu, Yixin Wu, Jun Sakuma
Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmf…
cs.CR2025
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
Rui Wen, Yiyong Liu, Michael Backes +1
Data reconstruction attacks, which aim to recover the training dataset of a target model with limited access, have gained increasing attention in recent years. However, there is cu…