most citedMixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

1 citations · 1 across the 1 of their papers we have counts for

collaborators

9 papers

cs.CL2026

Scaling Laws for Code: A More Data-Hungry Regime

Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu +5

Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantly analyzed on Natural Language (…

cs.CL20261 cited

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

Houyi Li, Ka Man Lo, Shijie Xuyang +7

Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense…

cs.CL2026

Is Compression Really Linear with Code Intelligence?

Shijie Xuyang, Xianzhen Luo, Zheng Chu +6

Understanding the relationship between data compression and the capabilities of Large Language Models (LLMs) is crucial, especially in specialized domains like code intelligence. P…

cs.CV2025

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

Tianhao Peng, Haochen Wang, Yuanxing Zhang +13

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understa…

cs.LG2025

Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Houyi Li, Wenzhen Zheng, Qiufeng Wang +10

The impressive capabilities of Large Language Models (LLMs) across diverse tasks are now well established, yet their effective deployment necessitates careful hyperparameter optimi…

cs.LG2025

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding

StepFun, :, Bin Wang +195

Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hard…