64 citations · 333 across the 16 of their papers we have counts for
6 papers · 1 filter
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan +11
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…
Improving Representational Continuity via Continued Pretraining
Michael Sun, Ananya Kumar, Divyam Madaan +1
We consider the continual representation learning setting: sequentially pretrain a model on tasks , and then adapt on a small amount of data from each t…
Calibrated ensembles can mitigate accuracy tradeoffs under distribution shift
Ananya Kumar, Tengyu Ma, Percy Liang +1
We often see undesirable tradeoffs in robust machine learning where out-of-distribution (OOD) accuracy is at odds with in-distribution (ID) accuracy: a robust classifier obtained v…
Extending the WILDS Benchmark for Unsupervised Adaptation
Shiori Sagawa, Pang Wei Koh, Tony Lee +17
Machine learning systems deployed in the wild are often trained on a source distribution but deployed on a different target distribution. Unlabeled data can be a powerful point of…
Convexified Convolutional Neural Networks
Yuchen Zhang, Percy Liang, Martin J. Wainwright
We describe the class of convexified convolutional neural networks (CCNNs), which capture the parameter sharing of convolutional neural networks in a convex manner. By representing…
Tensor Factorization via Matrix Factorization
Volodymyr Kuleshov, Arun Tejasvi Chaganty, Percy Liang
Tensor factorization arises in many machine learning applications, such knowledge base modeling and parameter estimation in latent variable models. However, numerical methods for t…