Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MESA: Improving MoE Safety Alignment via Decentralized Expertise
Yitong Sun, Yao Huang, Teng Li +5
Mixture-of-Experts (MoE) architectures scale Large Language Models (LLMs) efficiently, enabling greater capacity with reduced computational cost by dynamically routing inputs to re…
cs.LG2024
On the Adversarial Transferability of Generalized "Skip Connections"
Yisen Wang, Yichuan Mo, Dongxian Wu +3
Skip connection is an essential ingredient for modern deep models to be deeper and more powerful. Despite their huge success in normal scenarios (state-of-the-art classification pe…