3 papers
cs.LG2026
A theoretical model for task routing in mixture-of-expert transformers
Vinoth Nandakumar, Yongli Xiang, Yunzhi Yao +2
Mixture-of-experts (MoE) layers enable the scaling of transformer models while keeping the inference compute fixed. While task-expert specialization has been observed in empirical…
cs.LG2026
Visual-Guided Key-Token Regularization for Multimodal Large Language Model Unlearning
Chengyi Cai, Zesheng Ye, Peike Li +3
Unlearning in Multimodal Large Language Models (MLLMs) prevents the model from revealing private information when queried about target images. Existing MLLM unlearning methods larg…
cs.CV2025
Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation
Jianyuan Guo, Peike Li, Trevor Cohn
Sign Language Translation (SLT) aims to map sign language videos to spoken language text. A common approach relies on gloss annotations as an intermediate representation, decomposi…