1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.CL2025
Automating Steering for Safe Multimodal Large Language Models
Lyucheng Wu, Mengru Wang, Ziwen Xu +4
Recent progress in Multimodal Large Language Models (MLLMs) has unlocked powerful cross-modal reasoning abilities, but also raised new safety concerns, particularly when faced with…
cs.CL2025
LookAhead Tuning: Safer Language Models via Partial Answer Previews
Kangwei Liu, Mengru Wang, Yujie Luo +7
Fine-tuning enables large language models (LLMs) to adapt to specific domains, but often compromises their previously established safety alignment. To mitigate the degradation of m…
cs.CR2024★ 1 cited
PhishIntel: Toward Practical Deployment of Reference-Based Phishing Detection
Yuexin Li, Hiok Kuek Tan, Qiaoran Meng +6
Phishing is a critical cyber threat, exploiting deceptive tactics to compromise victims and cause significant financial losses. While reference-based phishing detectors (RBPDs) hav…