collaborators

11 papers

cs.AR2026

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

Jiahao Lin, Alish Kanani, Sangwan Lee +2

Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware…

cs.AR2026

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

Alish Kanani, Layan Badawi, Umit Y. Ogras

Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving…

cs.AR2026

Uni-SFU: Algorithm-HW Co-Design for Universal SFUs via Mixed-Degree Piecewise Approximation

Miao Sun, Yucheng Huang, Mingcong Cao +3

Nonlinear activation functions are essential to modern deep neural networks (DNNs), but their hardware evaluation places significant pressure on the special-function units (SFUs) o…

cs.AR2026

ReVolt: Power Delivery Network-Aware Voltage Droop Control for 2.5D PIM Chiplet Architectures

Vibhanshu Sharma, Alish Kanani, Miao Sun +3

Processing-in-memory (PIM)-based 2.5D multi-chiplet platforms are enablers for machine learning (ML) workloads. However, their performance is affected by the power delivery network…

cs.AR2025

CHIPSIM: A Co-Simulation Framework for Deep Learning on Chiplet-Based Systems

Lukas Pfromm, Alish Kanani, Harsh Sharma +3

Due to reduced manufacturing yields, traditional monolithic chips cannot keep up with the compute, memory, and communication demands of data-intensive applications, such as rapidly…

cs.AR2025

THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures

Alish Kanani, Lukas Pfromm, Harsh Sharma +3

Chiplet-based integration enables large-scale systems that combine diverse technologies, enabling higher yield, lower costs, and scalability, making them well-suited to AI workload…