3 papers
cs.CV2026
VkSplat: High-Performance 3DGS Training in Vulkan Compute
Jingxiang Chen, Mohamed Ibrahim, Yang Liu
We present VkSplat, a high-performance, cross-vendor 3D Gaussian Splatting (3DGS) training pipeline implemented fully in Vulkan compute, addressing performance and compatibility li…
cs.DC2026
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +3
Offloading communication to existing direct memory access (DMA) engines, available on most state-of-the-art commercial GPUs, has emerged as an interesting and low-cost solution to…
cs.AR2025
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
Varsha Singhania, Shaizeen Aga, Mohamed Assem Ibrahim
Ubiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power co…