activity
20182022
most citedUsing DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

299 citations · 302 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2022299 cited

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Shaden Smith, Mostofa Patwary, Brandon Norick +17

Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few…

cs.AI2020

Extracting Angina Symptoms from Clinical Notes Using Pre-Trained Transformer Architectures

Aaron S. Eisman, Nishant R. Shah, Carsten Eickhoff +4

Anginal symptoms can connote increased cardiac risk and a need for change in cardiovascular management. This study evaluated the potential to extract these symptoms from physician…

cs.LG2020

A Transformer-based Framework for Multivariate Time Series Representation Learning

George Zerveas, Srideepika Jayaraman, Dhaval Patel +2

In this work we propose for the first time a transformer-based framework for unsupervised representation learning of multivariate time series. Pre-trained models can be potentially…

cs.IR20203 cited

Brown University at TREC Deep Learning 2019

George Zerveas, Ruochen Zhang, Leila Kim +1

This paper describes Brown University's submission to the TREC 2019 Deep Learning track. We followed a 2-phase method for producing a ranking of passages for a given input query: I…

cs.LG2018

Improving Clinical Predictions through Unsupervised Time Series Representation Learning

Xinrui Lyu, Matthias Hueser, Stephanie L. Hyland +2

In this work, we investigate unsupervised representation learning on medical time series, which bears the promise of leveraging copious amounts of existing unlabeled data in order…