activity
20192022
most citedTutel: Adaptive Mixture-of-Experts at Scale

32 citations · 43 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2022★ 2 cited

An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable Context

Xiaoyu Chen, Xiangming Zhu, Yufeng Zheng +8

One of the key challenges in deploying RL to real-world applications is to adapt to variations of unknown environment contexts, such as changing terrains in robotic tasks and fluct…

cs.DC2022★ 32 cited

Tutel: Adaptive Mixture-of-Experts at Scale

Changho Hwang, Wei Cui, Yifan Xiong +12

Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…

cs.DC2021★ 1 cited

CrossoverScheduler: Overlapping Multiple Distributed Training Applications in a Crossover Manner

Cheng Luo, Lei Qu, Youshan Miao +2

Distributed deep learning workloads include throughput-intensive training tasks on the GPU clusters, where the Distributed Stochastic Gradient Descent (SGD) incurs significant comm…

cs.DC2020

Simulating Performance of ML Systems with Offline Profiling

Hongming Huang, Peng Cheng, Hong Xu +1

We advocate that simulation based on offline profiling is a promising approach to better understand and improve the complex ML systems. Our approach uses operation-level profiling…

cs.CR2019★ 4 cited

BotGraph: Web Bot Detection Based on Sitemap

Yang Luo, Guozhen She, Peng Cheng +1

The web bots have been blamed for consuming large amount of Internet traffic and undermining the interest of the scraped sites for years. Traditional bot detection studies focus ma…

cs.NI2019

NetKernel: Making Network Stack Part of the Virtualized Infrastructure

Zhixiong Niu, Hong Xu, Peng Cheng +4

This paper presents a system called NetKernel that decouples the network stack from the guest virtual machine and offers it as an independent module. NetKernel represents a new par…