collaborators

9 papers

cs.LG2026

MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core

Dennis Liu, Zijie Yan, Xin Yao +15

Mixture of Experts (MoE) models enhance neural network scalability by dynamically selecting relevant experts per input token, enabling larger model sizes while maintaining manageab…

cs.LG2026

Do Transformers Have the Ability for Periodicity Generalization?

Huanyu Liu, Ge Li, Yihong Dong +7

Large language models (LLMs) based on the Transformer have demonstrated strong performance across diverse tasks. However, current models still exhibit substantial limitations in ou…

cs.CL2025

SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression

Biao Zhang, Lixin Chen, Tong Liu +1

Large language models (LLMs) generate high-dimensional embeddings that capture rich semantic and syntactic information. However, high-dimensional embeddings exacerbate computationa…

cs.SE2025

FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

Hongda Zhu, Yiwen Zhang, Bing Zhao +6

Large Language Models (LLMs) have made significant strides in front-end code generation. However, existing benchmarks exhibit several critical limitations: many tasks are overly si…

cs.CL2025

Upcycling Large Language Models into Mixture of Experts

Ethan He, Abhinav Khattar, Ryan Prenger +7

Upcycling pre-trained dense language models into sparse mixture-of-experts (MoE) models is an efficient approach to increase the model capacity of already trained models. However,…

cs.CL2025

Taming the Titans: A Survey of Efficient LLM Inference Serving

Ranran Zhen, Juntao Li, Yixin Ji +7

Large Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applicat…