4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CR2025
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
Yuanbo Xie, Yingjie Zhang, Tianyun Liu +2
Jailbreak attacks pose persistent threats to large language models (LLMs). Current safety alignment methods have attempted to address these issues, but they experience two signific…
cs.MM2025★ 4 cited
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
Taoyu Su, Jiawei Sheng, Duohe Ma +5
Multi-Modal Entity Alignment (MMEA) aims to retrieve equivalent entities from different Multi-Modal Knowledge Graphs (MMKGs), a critical information retrieval task. Existing studie…