◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Joe Chau

3 papers

No researched profile yet.

papers

Publications (3)

cs.DC2023

Tutel: Adaptive Mixture-of-Experts at Scale

Changho Hwang, Wei Cui, Yifan Xiong +12

Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…

cs.DC2024

SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation

Yifan Xiong, Yuting Jiang, Ziyue Yang +17

Reliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently…

cs.LG2023

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17

In this paper, we explore FP8 low-bit data formats for efficient training of large language models (LLMs). Our key insight is that most variables, such as gradients and optimizer s…

◍wovepaper

A living map of arXiv — papers, researchers, institutions.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Sign in
  • Library
  • Chat
Data
  • arXiv.org
  • Latest RSS
Metadata from arXiv.org · Not affiliated with arXiv