◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Stephen Roller

21 papers hereh-index 2112.7k citations49 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author14
  • last author1

Across the 18 of 21 papers where every author was matched, so the position is known.

fields
  • cs.CL14
  • cs.LG5
  • cs.AI1
  • cs.CV1
same name
  • Stephen Roller — 3 papers, h 13
  • Stephen Roller — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20182025
most citedNeural Text Generation with Unlikelihood Training

241 citations · 465 across the 16 of their papers we have counts for

collaborators
Showing 2021Show all

4 papers · 1 filter

cs.CL2021★ 4 cited

Teaching Models new APIs: Domain-Agnostic Simulators for Task Oriented Dialogue

Moya Chen, Paul A. Crook, Stephen Roller

We demonstrate that large language models are able to simulate Task Oriented Dialogues in novel domains, provided only with an API implementation and a list of goals. We show these…

cs.LG2021★ 48 cited

Hash Layers For Large Sparse Models

Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam +1

We investigate the training of sparse layers that use different parameters for different inputs based on hashing in large Transformer models. Specifically, we modify the feedforwar…

cs.LG2021★ 5 cited

Not All Memories are Created Equal: Learning to Forget by Expiring

Sainbayar Sukhbaatar, Da Ju, Spencer Poff +4

Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory. Recent work investigated mechanisms to reduce the computational cost of…

cs.LG2021★ 3 cited

Staircase Attention for Recurrent Processing of Sequences

Da Ju, Stephen Roller, Sainbayar Sukhbaatar +1

Attention mechanisms have become a standard tool for sequence modeling tasks, in particular by stacking self-attention layers over the entire input sequence as in the Transformer a…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.