activity
20162020
most citedA Deep Reinforcement Learning Chatbot

200 citations · 318 across the 5 of their papers we have counts for

collaborators

11 papers

cs.CL20203 cited

Multi-scale Transformer Language Models

Sandeep Subramanian, Ronan Collobert, Marc'Aurelio Ranzato +1

We investigate multi-scale transformer language models that learn representations of text at multiple scales, and present three different architectures that have an inductive bias…

cs.CL2019

On Extractive and Abstractive Neural Document Summarization with Transformer Language Models

Sandeep Subramanian, Raymond Li, Jonathan Pilault +1

We present a method to produce abstractive summaries of long documents that exceed several thousand words via neural abstractive summarization. We perform a simple extractive step…

cs.LG2019

State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations

Alex Lamb, Jonathan Binas, Anirudh Goyal +5

Machine learning promises methods that generalize well from finite labeled data. However, the brittleness of existing neural net approaches is revealed by notable failures, such as…

cs.CL2018

Multiple-Attribute Text Style Transfer

Sandeep Subramanian, Guillaume Lample, Eric Michael Smith +3

The dominant approach to unsupervised "style transfer" in text is based on the idea of learning a latent representation, which is independent of the attributes specifying its "styl…

stat.ML2018

Fortified Networks: Improving the Robustness of Deep Networks by Modeling the Manifold of Hidden Representations

Alex Lamb, Jonathan Binas, Anirudh Goyal +4

Deep networks have achieved impressive results across a variety of important tasks. However a known weakness is a failure to perform well when evaluated on data which differ from t…

cs.CL2018

Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning

Sandeep Subramanian, Adam Trischler, Yoshua Bengio +1

A lot of the recent success in natural language processing (NLP) has been driven by distributed vector representations of words trained on large amounts of text in an unsupervised…