2 papers
cs.AR2026
Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis
Zhongchun Zhou, Yuhang Gu, Chengtao Lai +4
To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization. These features…
cs.AR2025
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
Zhongchun Zhou, Chengtao Lai, Yuhang Gu +1
The rapid adoption of large language models (LLMs) is pushing AI accelerators toward increasingly powerful and specialized designs. Instead of further complicating software develop…