2 papers
cs.AR2026
Technology solutions targeting the performance of gen-AI inference in resource constrained platforms
Joyjit Kundu, Joshua Klein, Aakash Patel +1
The rise of generative AI workloads, particularly language model inference, is intensifying on/off-chip memory pressure. Multimodal inputs such as video streams or images and downs…
cs.ET2025
LionHeart: A Layer-based Mapping Framework for Heterogeneous Systems with Analog In-Memory Computing Tiles
Corey Lammie, Yuxuan Wang, Flavio Ponzina +7
When arranged in a crossbar configuration, resistive memory devices can be used to execute Matrix-Vector Multiplications (MVMs), the most dominant operation of many Machine Learnin…