318 citations
- Amazon (Germany)DE1 paper
- Duke UniversityUS1 paper
- Karlsruhe Institute of TechnologyDE1 paper
- Martin Luther University Halle-WittenbergDE1 paper
- Microsoft (Finland)FI1 paper
- Microsoft Research (United Kingdom)GB1 paper
- Princeton UniversityUS1 paper
- University of California, Santa BarbaraUS1 paper
- University of FreiburgDE1 paper
- University of Illinois Urbana-ChampaignUS1 paper
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2020★ 20 cited
AdaScale SGD: A User-Friendly Algorithm for Distributed Training
Tyler B. Johnson, Pulkit Agrawal, Haijie Gu +1
When using large-batch training to speed up stochastic gradient descent, learning rates must adapt to new batch sizes in order to maximize speed-ups and preserve model quality. Re-…
cs.LG2020★ 42 cited
Learning to Branch for Multi-Task Learning
Pengsheng Guo, Chen-Yu Lee, Daniel Ulbricht
Training multiple tasks jointly in one deep network yields reduced latency during inference and better performance over the single-task counterpart by sharing certain layers of a n…
cs.LG2020★ 8 cited
Privacy-preserving Learning via Deep Net Pruning
Yangsibo Huang, Yushan Su, Sachin Ravi +3
This paper attempts to answer the question whether neural network pruning can be used as a tool to achieve differential privacy without losing much data utility. As a first step to…