12 papers
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
Prabhu Vellaisamy, Shreesh Tripathi, Vignesh Natarajan +3
Large Language Model (LLM) inference is widely used in interactive assistants and agentic systems. In latency-sensitive deployments, inference time can become dominated by host-sid…
A-Graph: A Unified Graph Representation for At-Will Simulation across System Stacks
Daniel Price, Prabhu Vellaisamy, Patricia Gonzalez +3
As computer systems continue to diversify across technologies, architectures, applications, and beyond, the relevant design space has become larger and more complex. Given such tre…
Mugi: Value Level Parallelism For Efficient LLMs
Daniel Price, Prabhu Vellaisamy, John Shen +1
Value level parallelism (VLP) has been proposed to improve the efficiency of large-batch, low-precision general matrix multiply (GEMM) between symmetric activations and weights. In…
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
Shanmuga Venkatachalam, Prabhu Vellaisamy, Harideep Nair +5
Leading experts from both communities have suggested the need to (re)connect research in neuroscience and artificial intelligence (AI) to accelerate the development of next-generat…
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
Prabhu Vellaisamy, Harideep Nair, Di Wu +2
General matrix multiplication (GEMM) is a fundamental operation in deep learning (DL). With DL moving increasingly toward low precision, recent works have proposed novel unary GEMM…
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
Devon Lister, Prabhu Vellaisamy, John Paul Shen +1
Temporal neural networks (TNNs) are neuromorphic neural networks that utilize bit-serial temporal coding. TNNs are composed of columns, which in turn employ neurons as their buildi…