3 papers
cs.CR2026
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Yi Shi, Tanyu Chen, Kai Shen
Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no…
cs.CL2025
Enhancing Meme Emotion Understanding with Multi-Level Modality Enhancement and Dual-Stage Modal Fusion
Yi Shi, Wenlong Meng, Zhenyuan Guo +2
With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion…
cs.CL2025
Be Cautious When Merging Unfamiliar LLMs: A Phishing Model Capable of Stealing Privacy
Zhenyuan Guo, Yi Shi, Wenlong Meng +3
Model merging is a widespread technology in large language models (LLMs) that integrates multiple task-specific LLMs into a unified one, enabling the merged model to inherit the sp…