3 papers
cs.CR2026
When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models
Ruihan Hu, Yu-Ming Shang, Wei Luo +2
Large Reasoning Models (LRMs) have rapidly gained prominence for their strong performance in solving complex tasks. Many modern black-box LRMs expose the intermediate reasoning tra…
cs.CR2026
From static to adaptive: immune memory-based jailbreak detection for large language models
Jun Leng, Yu Liu, Litian Zhang +3
Large Language Models (LLMs) serve as the backbone of modern AI systems, yet they remain susceptible to adversarial jailbreak attacks. Consequently, robust detection of such malici…
cs.CL2025
Automated Detection of Pre-training Text in Black-box LLMs
Ruihan Hu, Yu-Ming Shang, Jiankun Peng +3
Detecting whether a given text is a member of the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing met…