activity
20242026
most citedPrediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

ESPO: Early-Stopping Proximal Policy Optimization

Zihang Li, Rui Zhou, Yingcheng Shi +8

When a large language model under reinforcement learning commits a wrong reasoning step early in a trajectory, standard algorithms force it to keep generating until the maximum hor…

cs.LG2025

Rank Also Matters: Hierarchical Configuration for Mixture of Adapter Experts in LLM Fine-Tuning

Peizhuang Cong, Wenpu Liu, Wenhan Yu +2

Large language models (LLMs) have demonstrated remarkable success across various tasks, accompanied by a continuous increase in their parameter size. Parameter-efficient fine-tunin…

cs.LG2024

BATON: Enhancing Batch-wise Inference Efficiency for Large Language Models via Dynamic Re-batching

Peizhuang Cong, Qizhi Chen, Haochen Zhao +1

The advanced capabilities of Large Language Models (LLMs) have inspired the development of various interactive web services or applications, such as ChatGPT, which offer query infe…

cs.LG2024

INT-FlashAttention: Enabling Flash Attention for INT8 Quantization

Shimao Chen, Zirui Liu, Zhiying Wu +6

As the foundation of large language models (LLMs), self-attention module faces the challenge of quadratic time and memory complexity with respect to sequence length. FlashAttention…

cs.LG20241 cited

Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing

Peizhuang Cong, Aomufei Yuan, Shimao Chen +3

MoE facilitates the development of large models by making the computational complexity of the model no longer scale linearly with increasing parameters. The learning sparse gating…