Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
The xPU-athalon: Quantifying the Competition of AI Acceleration
Alicia Golden, Carole-Jean Wu, Gu-Yeon Wei +1
The push for greater efficiency in AI computation has given rise to an array of accelerator architectures that increasingly challenge the GPU's long-standing dominance. In this wor…
cs.AR2026
RPU -- A Reasoning Processing Unit
Matthew Adiletta, Gu-Yeon Wei, David Brooks
Large language model (LLM) inference performance is increasingly bottlenecked by the memory wall. While GPUs continue to scale raw compute throughput, they struggle to deliver scal…
cs.AR2024
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
Yun-Chen Lo, Gu-Yeon Wei, David Brooks
As cutting-edge large language models (LLMs) continue to transform various industries, their fast-growing model size and sequence length have led to memory traffic and capacity cha…