8 citations · 9 across the 6 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
Chuhan Zhang, Ye Zhang, Bowen Shi +5
Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechan…
cs.CR2024
CAMH: Advancing Model Hijacking Attack in Machine Learning
Xing He, Jiahao Chen, Yuwen Pu +5
In the burgeoning domain of machine learning, the reliance on third-party services for model training and the adoption of pre-trained models have surged. However, this reliance int…