activity
20212026
most citedTo Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis

20 citations · 82 across the 18 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.LG2021★ 9 cited

Large-Scale Deep Learning Optimizations: A Comprehensive Survey

Xiaoxin He, Fuzhao Xue, Xiaozhe Ren +1

Deep learning have achieved promising results on a wide spectrum of AI applications. Larger datasets and models consistently yield better performance. However, we generally spend l…

cs.LG2021★ 4 cited

Cross-token Modeling with Conditional Computation

Yuxuan Lou, Fuzhao Xue, Zangwei Zheng +1

Mixture-of-Experts (MoE), a conditional computation architecture, achieved promising performance by scaling local module (i.e. feed-forward network) of transformer. However, scalin…

cs.CL2021

Automated Audio Captioning using Transfer Learning and Reconstruction Latent Space Similarity Regularization

Andrew Koh, Fuzhao Xue, Eng Siong Chng

In this paper, we examine the use of Transfer Learning using Pretrained Audio Neural Networks (PANNs), and propose an architecture that is able to better leverage the acoustic feat…

cs.LG2021★ 3 cited

Go Wider Instead of Deeper

Fuzhao Xue, Ziji Shi, Futao Wei +3

More transformer blocks with residual connections have recently achieved impressive results on various tasks. To achieve better performance with fewer trainable parameters, recent…

cs.LG2021★ 4 cited

Sequence Parallelism: Long Sequence Training from System Perspective

Shenggui Li, Fuzhao Xue, Chaitanya Baranwal +2

Transformer achieves promising results on various tasks. However, self-attention suffers from quadratic memory requirements with respect to the sequence length. Existing work focus…