Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis
Zhongchun Zhou, Yuhang Gu, Chengtao Lai +4
To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization. These features…
cs.AR2025
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
Jinhao Li, Jiaming Xu, Shan Huang +9
Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation. Compared to non-generative LLM…
cs.AR2024
MARCA: Mamba Accelerator with ReConfigurable Architecture
Jinhao Li, Shan Huang, Jiaming Xu +4
We propose a Mamba accelerator with reconfigurable architecture, MARCA.We propose three novel approaches in this paper. (1) Reduction alternative PE array architecture for both lin…