2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2025
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Rishabh Tiwari, Haocheng Xi, Aditya Tomar +7
Large Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference. In th…
cs.CL2024★ 2 cited
Characterizing Prompt Compression Methods for Long Context Inference
Siddharth Jha, Lutfi Eren Erdogan, Sehoon Kim +2
Long context inference presents challenges at the system level with increased compute and memory requirements, as well as from an accuracy perspective in being able to reason over…
cs.LG2024
Reliable edge machine learning hardware for scientific applications
Tommaso Baldi, Javier Campos, Ben Hawks +15
Extreme data rate scientific experiments create massive amounts of data that require efficient ML edge processing. This leads to unique validation challenges for VLSI implementatio…