3 papers
cs.AI2025
Pay More Attention to the Robustness of Prompt for Instruction Data Mining
Qiang Wang, Dawei Feng, Xu Zhang +4
Instruction tuning has emerged as a paramount method for tailoring the behaviors of LLMs. Recent work has unveiled the potential for LLMs to achieve high performance through fine-t…
cs.DC2024
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
Zhigang Wang, Xu Zhang, Ning Wang +5
Transformer-based models are becoming deeper and larger recently. For better scalability, an underlying training solution in industry is to split billions of parameters (tensors) i…
cs.AR2021
Asynchronous Memory Access Unit for General Purpose Processors
Luming Wang, Xu Zhang, Tianyue Lu +1
In future data centers, applications will make heavy use of far memory (including disaggregated memory pools and NVM). The access latency of far memory is more widely distributed t…