◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

M. Andersch

2 papers hereh-index 12897 citations23 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG1
  • cs.NE1

identity via Semantic Scholar / OpenAlex

most citedReducing Activation Recomputation in Large Transformer Models

55 citations · 55 across the 1 of their papers we have counts for

collaborators

2 papers

cs.LG2022★ 55 cited

Reducing Activation Recomputation in Large Transformer Models

Vijay Korthikanti, Jared Casper, Sangkug Lym +4

Training large transformer models is one of the most important computational challenges of modern AI. In this paper, we show how to significantly accelerate training of large trans…

cs.NE2018

Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip

Feiwen Zhu, Jeff Pool, Michael Andersch +2

Recurrent Neural Networks (RNNs) are powerful tools for solving sequence-based problems, but their efficacy and execution time are dependent on the size of the network. Following r…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.