5 papers
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
Jingcong Liang, Shijun Wan, Xuehai Wu +5
Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints.…
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
Jingcong Liang, Siyuan Wang, Miren Tian +3
Mixture-of-Experts (MoE) enables efficient scaling of large language models (LLMs) with sparsely activated experts during inference. To effectively deploy large MoE models on memor…
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
Dianyi Wang, Siyuan Wang, Zejun Li +6
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across multi-modal tasks by scaling model size and training data. However, these dense LVLMs incur sig…
Mixture of Lookup Experts
Shibo Jie, Yehui Tang, Kai Han +4
Mixture-of-Experts (MoE) activates only a subset of experts during inference, allowing the model to maintain low inference FLOPs and latency even as the parameter count scales up.…
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
Haoyu Wang, Tong Teng, Tianyu Guo +4
Handling long-context sequences efficiently remains a significant challenge in large language models (LLMs). Existing methods for token selection in sequence extrapolation either e…