activity
20192026
most citedAIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.ARShow all

6 papers · 1 filter

cs.AR2026

HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

Haochen Huang, Shuzhang Zhong, Shengxuan Qiu +8

Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However,…

cs.AR2026

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation

Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6

Deploying Video Diffusion Models (VDMs) on edge devices is appealing for localized and privacy-preserving generation, but their iterative Transformer-based denoising remains too sl…

cs.AR2026

A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators

Cong Li, Chenhao Xue, Yi Ren +11

Large language models (LLMs) exhibit memory-intensive behavior during decoding, making it a key bottleneck in LLM inference. To accelerate decoding execution, hybrid-bonding-based…

cs.AR2026

Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator

Cong Li, Yihan Yin, Chenhao Xue +7

Large language models (LLMs) have been widely deployed for online generative services, where numerous LLM instances jointly handle workloads with fluctuating request arrival rates…

cs.AR20251 cited

AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM

Yuanpeng Zhang, Xing Hu, Xi Chen +10

SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computation…

cs.AR2025

Enabling Efficient Transaction Processing on CXL-Based Memory Sharing

Zhao Wang, Yiqi Chen, Cong Li +5

Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute…