4 papers
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
Hanzhi Zhang, Qiao Zhang, Qinglei Cao +4
Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing me…
PHASE: Physics-Integrated, Heterogeneity-Aware Surrogates for Scientific Simulations
Dawei Gao, Dali Wang, Zhuowei Gu +5
Large-scale numerical simulations underpin modern scientific discovery but remain constrained by prohibitive computational costs. AI surrogates offer acceleration, yet adoption in…
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
Qiao Zhang, Rabab Alomairy, Dali Wang +2
General Matrix Multiplication (GEMM) is a critical operation underpinning a wide range of applications in high-performance computing (HPC) and artificial intelligence (AI). The eme…
Kilometer-Scale E3SM Land Model Simulation over North America
Dali Wang, Chen Wang, Qinglei Cao +9
The development of a kilometer-scale E3SM Land Model (km-scale ELM) is an integral part of the E3SM project, which seeks to advance energy-related Earth system science research wit…