◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Deyu Fu

2 papers hereh-index 4126 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1

Across the 1 of 2 papers where every author was matched, so the position is known.

fields
  • cs.DC1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

2 papers

cs.LG2026

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Mikail Khona, Aditya Vavre, Boxiang Wang +11

Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. I…

cs.DC2026

Scalable Training of Mixture-of-Experts Models with Megatron Core

Zijie Yan, Hongxiao Bai, Xin Yao +42

Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total pa…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.