activity
20172024
most citedFoundation Transformers

13 citations · 35 across the 8 of their papers we have counts for

collaborators

9 papers

cs.LG20226 cited

TorchScale: Transformers at Scale

Shuming Ma, Hongyu Wang, Shaohan Huang +8

Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…

cs.CL20223 cited

Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning

Barun Patra, Saksham Singhal, Shaohan Huang +5

In this paper, we elaborate upon recipes for building multilingual representation models that are not only competitive with existing state-of-the-art models but are also more param…

cs.LG202213 cited

Foundation Transformers

Hongyu Wang, Shuming Ma, Shaohan Huang +12

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…

cs.CL20221 cited

On Efficiently Acquiring Annotations for Multilingual Models

Joel Ruben Antony Moniz, Barun Patra, Matthew R. Gormley

When tasked with supporting multiple languages for a given problem, two approaches have arisen: training a model for each language with the annotation budget divided equally among…

cs.CL2020

To Schedule or not to Schedule: Extracting Task Specific Temporal Entities and Associated Negation Constraints

Barun Patra, Chala Fufa, Pamela Bhattacharya +1

State of the art research for date-time entity extraction from text is task agnostic. Consequently, while the methods proposed in literature perform well for generic date-time extr…

cs.CL2020

ScopeIt: Scoping Task Relevant Sentences in Documents

Vishwas Suryanarayanan, Barun Patra, Pamela Bhattacharya +2

Intelligent assistants like Cortana, Siri, Alexa, and Google Assistant are trained to parse information when the conversation is synchronous and short; however, for email-based con…