1 paper
Yiqi Liu, Yudong Pan, Mengdi Wang +5
Conventional LLM inference architectures suffer from high energy and latency due to frequent data movement across memory hierarchies. We propose Ouroboros, a wafer-scale SRAM-based…