◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Chandler Zhou

3 papers hereh-index 335 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author1

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.DC1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

activity
20232026
collaborators

3 papers

cs.DC2026

Scalable Training of Mixture-of-Experts Models with Megatron Core

Zijie Yan, Hongxiao Bai, Xin Yao +42

Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total pa…

cs.LG2025

MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core

Dennis Liu, Zijie Yan, Xin Yao +15

Mixture of Experts (MoE) models enhance neural network scalability by dynamically selecting relevant experts per input token, enabling larger model sizes while maintaining manageab…

cs.CL2023

Aligning Language Models with Offline Learning from Human Feedback

Jian Hu, Li Tao, June Yang +1

Learning from human preferences is crucial for language models (LMs) to effectively cater to human needs and societal values. Previous research has made notable progress by leverag…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.