most citedMixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

2 citations · 3 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix Perspective

Xianzhen Luo, Jinyang Huang, Wenzhen Zheng +5

Evaluating test cases automatically generated by Large Language Models (LLMs) is a critical yet challenging task. Existing benchmarks often evaluate the exclusion ratio on large, u…

cs.CL2025

Scaling Laws for Code: A More Data-Hungry Regime

Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu +5

Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantly analyzed on Natural Language (…

cs.LG2025

Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models

Houyi Li, Wenzhen Zheng, Qiufeng Wang +8

Training Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-int…

cs.CL2025★ 2 cited

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

Houyi Li, Ka Man Lo, Shijie Xuyang +7

Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense…

cs.LG2025★ 1 cited

Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Houyi Li, Wenzhen Zheng, Qiufeng Wang +10

The impressive capabilities of Large Language Models (LLMs) across diverse tasks are now well established, yet their effective deployment necessitates careful hyperparameter optimi…

cs.CL2024

Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale

Wenzhen Zheng, Wenbo Pan, Xu Xu +3

In recent years, Large Language Models (LLMs) have made significant strides towards Artificial General Intelligence. However, training these models from scratch requires substantia…