17 papers
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
Robin Geens, Jonas De Schouwer, Marian Verhelst +1
The Hardware Lottery posits that research directions are dictated by available silicon compute platforms. We identify a derivative phenomenon, the Hyperscale Lottery, where model a…
Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
Yasmine Omri, Ziyu Gan, Zachary Broveak +6
LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories. Realizing this at scale requires agents to persistentl…
SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers
Wonsuk Jang, Thierry Tambe
Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder edge deployment. Quantization can…
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
Yuzong Chen, Chao Fang, Xilai Dai +4
The substantial memory bandwidth and computational demands of large language models (LLMs) present critical challenges for efficient inference. To tackle this, the literature has e…
AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents
Hojoon Kim, Yuheng Wu, Thierry Tambe
Embodied AI agents increasingly rely on large language models (LLMs) for planning, yet per-step LLM calls impose severe latency and cost. In this paper, we show that embodied tasks…
Heterogeneous Memory Design Exploration for AI Accelerators with a Gain Cell Memory Compiler
Xinxin Wang, Lixian Yan, Shuhan Liu +10
As memory increasingly dominates system cost and energy, heterogeneous on-chip memory systems that combine technologies with complementary characteristics are becoming essential. G…