activity
20242026
collaborators

5 papers

cs.AR2026

Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy

Peichen Xie, Shuotao Xu, Yang Wang +2

Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose…

cs.AR2026

LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis

Tao Zhang, Rui Ma, Shuotao Xu +2

GPU design space exploration (DSE) for modern AI workloads, such as Large-Language Model (LLM) inference, is challenging because of GPUs' vast, multi-modal design spaces, high simu…

cs.DC2025

FengHuang: Next-Generation Memory Orchestration for AI Inferencing

Jiamin Li, Lei Qu, Tao Zhang +4

This document presents a vision for a novel AI infrastructure design that has been initially validated through inference simulations on state-of-the-art large language models. Adva…

cs.IR2024

SPFresh: Incremental In-Place Update for Billion-Scale Vector Search

Yuming Xu, Hengyu Liang, Jin Li +9

Approximate Nearest Neighbor Search (ANNS) is now widely used in various applications, ranging from information retrieval, question answering, and recommendation, to search for sim…

cs.AR2024

NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering

Zhe Zhou, Yiqi Chen, Tao Zhang +8

The Compute Express Link (CXL) interconnect makes it feasible to integrate diverse types of memory into servers via its byte-addressable SerDes links. Considering the various acces…