activity
20242026
collaborators

5 papers

cs.AR2026

The xPU-athalon: Quantifying the Competition of AI Acceleration

Alicia Golden, Carole-Jean Wu, Gu-Yeon Wei +1

The push for greater efficiency in AI computation has given rise to an array of accelerator architectures that increasingly challenge the GPU's long-standing dominance. In this wor…

cs.AR2026

RPU -- A Reasoning Processing Unit

Matthew Adiletta, Gu-Yeon Wei, David Brooks

Large language model (LLM) inference performance is increasingly bottlenecked by the memory wall. While GPUs continue to scale raw compute throughput, they struggle to deliver scal…

cs.LG2025

The Energy Cost of Reasoning: Analyzing Energy Usage in LLMs with Test-time Compute

Yunho Jin, Gu-Yeon Wei, David Brooks

Scaling large language models (LLMs) has driven significant advancements, yet it faces diminishing returns and escalating energy demands. This work explores how test-time compute (…

cs.AI2025

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices

Yuji Chai, Mujin Kwen, David Brooks +1

Deploying LLMs on edge devices presents serious technical challenges. Memory elasticity is crucial for edge devices with unified memory, where memory is shared and fluctuates dynam…

cs.AR2024

Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models

Yun-Chen Lo, Gu-Yeon Wei, David Brooks

As cutting-edge large language models (LLMs) continue to transform various industries, their fast-growing model size and sequence length have led to memory traffic and capacity cha…