5 papers
TraceNAS: Zero-shot LLM Pruning via Gradient Trace Correlation
Prajna G. Malettira, Manish Nagaraj, Arjun Roy +2
Structured pruning is essential for efficient deployment of Large Language Models (LLMs). The varying sensitivity of LLM sub-blocks to pruning necessitates the identification of op…
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
Shubham Negi, Kaushik Roy
The rapid adoption of Large Language Models (LLMs) has driven a growing demand for efficient inference, particularly in latency-sensitive applications such as chatbots and personal…
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
Shubham Negi, Manik Singhal, Aayush Ankit +2
Modern machine learning accelerators are designed to efficiently execute deep neural networks (DNNs) by optimizing data movement, memory hierarchy, and compute throughput. However,…
TSkips: Efficiency Through Explicit Temporal Delay Connections in Spiking Neural Networks
Prajna G. Malettira, Shubham Negi, Wachirawit Ponghiran +1
Spiking Neural Networks (SNNs) with their bio-inspired Leaky Integrate-and-Fire (LIF) neurons inherently capture temporal information. This makes them well-suited for sequential ta…
SpiDR: A Reconfigurable Digital Compute-in-Memory Spiking Neural Network Accelerator for Event-based Perception
Deepika Sharma, Shubham Negi, Trishit Dutta +2
Spiking Neural Networks (SNNs), with their inherent recurrence, offer an efficient method for processing the asynchronous temporal data generated by Dynamic Vision Sensors (DVS), m…