2 papers
cs.LG2025
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
Pouya Hamadanian, Sadjad Fouladi
We introduce Glinthawk, an architecture for offline Large Language Model (LLM) inference. By leveraging a two-tiered structure, Glinthawk optimizes the utilization of the high-end…
cs.LG2020
Real-Time Video Inference on Edge Devices via Adaptive Model Streaming
Mehrdad Khani, Pouya Hamadanian, Arash Nasr-Esfahany +1
Real-time video inference on edge devices like mobile phones and drones is challenging due to the high computation cost of Deep Neural Networks. We present Adaptive Model Streaming…