51 citations · 152 across the 6 of their papers we have counts for
5 papers · 1 filter
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
Zichang Liu, Jue Wang, Tri Dao +8
Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…
Auto-Differentiation of Relational Computations for Very Large Scale Machine Learning
Yuxin Tang, Zhimin Ding, Dimitrije Jankov +3
The relational data model was designed to facilitate large-scale data management and analytics. We consider the problem of how to differentiate computations expressed relationally.…
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan +11
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…
A Federated Learning Framework for Healthcare IoT devices
Binhang Yuan, Song Ge, Wenhui Xing
The Internet of Things (IoT) revolution has shown potential to give rise to many medical applications with access to large volumes of healthcare data collected by IoT devices. Howe…
WaveletAE: A Wavelet-enhanced Autoencoder for Wind Turbine Blade Icing Detection
Binhang Yuan, Chen Wang, Chen Luo +4
Wind power, as an alternative to burning fossil fuels, is abundant and inexhaustible. To fully utilize wind power, wind farms are usually located in areas of high altitude and faci…