13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 13 cited
Embracing Federated Learning: Enabling Weak Client Participation via Partial Model Training
Sunwoo Lee, Tuo Zhang, Saurav Prakash +2
In Federated Learning (FL), clients may have weak devices that cannot train the full model or even hold it in their memory space. To implement large-scale FL applications, thus, it…
cs.LG2024
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
Yue Niu, Saurav Prakash, Salman Avestimehr
We propose a new attention mechanism with linear complexity, ATP, that fixates \textbf{A}ttention on \textbf{T}op \textbf{P}rincipal keys, rather than on each individual token. Par…