5 papers
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models
Fan Mo, Yuxuan Han, Geng Zhang +2
Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse ac…
MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents
Yv Zhang, Hao Sun, Hao Fang +5
External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences. However, this paradigm introduces a cri…
END: Early Noise Dropping for Efficient and Effective Context Denoising
Hongye Jin, Pei Chen, Jingfeng Yang +11
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, they are often distracted by irrelevant or…
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
Shuoyang Sun, Chang Dai, Hao Fang +6
Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens and verifying them with a tar…
Enhancing Efficiency in Multidevice Federated Learning through Data Selection
Fan Mo, Mohammad Malekzadeh, Soumyajit Chatterjee +2
Ubiquitous wearable and mobile devices provide access to a diverse set of data. However, the mobility demand for our devices naturally imposes constraints on their computational an…