5 papers
ADEPT: Architecture-Driven Energy-Efficient CNN Fine-Tuning on PIM Accelerators
Pratyush Dhingra, Vibhanshu Sharma, Janardhan Rao Doppa +1
Processing-in-memory-based (PIM) architectures have emerged as a promising solution for accelerating Convolutional Neural Network (CNN) workloads at the edge. Fine-tuning pre-train…
ThRIve: Thermally Robust CNN Inference via Low-Rank Adaptation in Heterogeneous PIM Architectures
Vibhanshu Sharma, Pratyush Dhingra, Janardhan Rao Doppa +1
Processing-In-Memory (PIM) has emerged as a promising technology for accelerating machine learning (ML) workloads. Specifically, non-volatile memory-based PIM architectures have en…
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa +1
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained…
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
Harsh Sharma, Pratyush Dhingra, Janardhan Rao Doppa +2
Transformers have revolutionized deep learning and generative modeling, enabling advancements in natural language processing tasks. However, the size of transformer models is incre…
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
Pratyush Dhingra, Janardhan Rao Doppa, Partha Pratim Pande
Transformer architectures have become the standard neural network model for various machine learning applications including natural language processing and computer vision. However…