collaborators

5 papers

cs.AR2026

M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization

Weiming Hu, Zihan Zhang, Haoyan Zhang +8

Existing low-bit Microscaling (MX) formats, such as MXFP4, often suffer from substantial accuracy degradation due to the use of a shared scaling factor with the Power-of-Two format…

cs.DC2025

Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage

Ziqi Yuan, Haoyang Zhang, Yirui Eric Zhou +4

We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our frame…

cs.AR2025

M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type

Weiming Hu, Haoyan Zhang, Cong Guo +7

Large language models (LLMs) are one of the most important killer computer applications. The recent algorithmic advancement proposes a fine-grained group-wise quantization for LLMs…

cs.AR2025

SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design

Haoyang Zhang, Yuqi Xue, Yirui Eric Zhou +2

The CXL-based solid-state drive (CXL-SSD) provides a promising approach towards scaling the main memory capacity at low cost. However, the CXL-SSD faces performance challenges due…

cs.AR202320 cited

G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations

Haoyang Zhang, Yirui Eric Zhou, Yuqi Xue +2

To break the GPU memory wall for scaling deep learning workloads, a variety of architecture and system techniques have been proposed recently. Their typical approaches include memo…