activity
20202024
most citedDeja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

19 citations · 48 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2024★ 1 cited

Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

Zichang Liu, Qingyun Liu, Yuening Li +6

Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality. For example, in our experimen…

cs.LG2023

Heterogeneous federated collaborative filtering using FAIR: Federated Averaging in Random Subspaces

Aditya Desai, Benjamin Meisburger, Zichang Liu +1

Recommendation systems (RS) for items (e.g., movies, books) and ads are widely used to tailor content to users on various internet platforms. Traditionally, recommendation models a…

cs.LG2023★ 19 cited

Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Zichang Liu, Jue Wang, Tri Dao +8

Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…

cs.LG2023★ 11 cited

Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time

Zichang Liu, Aditya Desai, Fangshuo Liao +5

Large language models(LLMs) have sparked a new wave of exciting AI applications. Hosting these models at scale requires significant memory resources. One crucial memory bottleneck…

cs.LG2022★ 9 cited

Learning Multimodal Data Augmentation in Feature Space

Zichang Liu, Zhiqiang Tang, Xingjian Shi +4

The ability to jointly learn from multiple modalities, such as text, audio, and visual data, is a defining feature of intelligent systems. While there have been promising advances…

cs.LG2021★ 1 cited

Efficient Inference via Universal LSH Kernel

Zichang Liu, Benjamin Coleman, Anshumali Shrivastava

Large machine learning models achieve unprecedented performance on various tasks and have evolved as the go-to technique. However, deploying these compute and memory hungry models…