3 papers
cs.AR2025
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
Zhongling Su, Rong Fu, Weihan Cao +4
Current FP8 grouped GEMM implementations require padding each group to a fixed alignment (e.g., 128), incurring memory and computational overhead. We propose \textit{TMA-Adaptive F…
cs.DC2025
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
Ding Tang, Jiecheng Zhou, Jiakai Hu +5
Recent advancements in large language models (LLMs) necessitate extensive computational resources, prompting the use of diverse hardware accelerators from multiple vendors. However…
cs.GR2025
TC-GS: A Faster Gaussian Splatting Module Utilizing Tensor Cores
Zimu Liao, Jifeng Ding, Siwei Cui +7
3D Gaussian Splatting (3DGS) renders pixels by rasterizing Gaussian primitives, where conditional alpha-blending dominates the computational cost in the rendering pipeline. This pa…