3 papers
cs.CL2026
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
Chenchen Tan, Youyang Qu, Xinghao Li +4
The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Along with this, the potential reten…
cs.CV2026
Causal Fingerprints of AI Generative Models
Hui Xu, Chi Liu, Congcong Zhu +3
AI generative models leave implicit traces in their generated images, which are commonly referred to as model fingerprints and are exploited for source attribution. Prior methods r…
cs.CL2025
When Harmless Words Harm: A New Threat to LLM Safety via Conceptual Triggers
Zhaoxin Zhang, Borui Chen, Yiming Hu +3
Recent research on large language model (LLM) jailbreaks has primarily focused on techniques that bypass safety mechanisms to elicit overtly harmful outputs. However, such efforts…