activity
20242026
collaborators

5 papers

cs.DC2026

FaCTz: Fast Critical-Point and Topology-Aware GPU Compression for Scientific Vector Fields

Mingze Xia, Yuxiao Li, Sheng Di +6

Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specifie…

cs.DC2025

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Daoce Wang, Pascal Grosset, Jesus Pulido +9

Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access d…

cs.AR2024

GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors

Chengming Zhang, Xinheng Ding, Baixi Sun +4

Heterogeneous hardware like Gaudi processor has been developed to enhance computations, especially matrix operations for Transformer-based large language models (LLMs) for generati…

cs.LG2024

SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training

Jinda Jia, Cong Xie, Hanlin Lu +8

Recent years have witnessed a clear trend towards language models with an ever-increasing number of parameters, as well as the growing training overhead and memory usage. Distribut…

cs.LG2024

FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources

Xiyuan Wei, Fanjiang Ye, Ori Yonay +4

Existing studies of training state-of-the-art Contrastive Language-Image Pretraining (CLIP) models on large-scale data involve hundreds of or even thousands of GPUs due to the requ…