26 citations · 26 across the 2 of their papers we have counts for
5 papers
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
Mert Hidayetoglu, Aurick Qiao, Michael Wyatt +3
Efficient parallelism is necessary for achieving low-latency, high-throughput inference with large language models (LLMs). Tensor parallelism (TP) is the state-of-the-art method fo…
Task-Based Programming for Adaptive Mesh Refinement in Compressible Flow Simulations
Anjiang Wei, Hang Song, Mert Hidayetoglu +3
High-order solvers for compressible flows are vital in scientific applications. Adaptive mesh refinement (AMR) is a key technique for reducing computational cost by concentrating r…
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
Samyam Rajbhandari, Mert Hidayetoglu, Aurick Qiao +5
Inference is now the dominant AI workload, yet existing systems force trade-offs between latency, throughput, and cost. Arctic Inference, an open-source vLLM plugin from Snowflake…
Petascale XCT: 3D Image Reconstruction with Hierarchical Communications on Multi-GPU Nodes
Mert Hidayetoglu, Tekin Bicer, Simon Garcia de Gonzalo +6
X-ray computed tomography is a commonly used technique for noninvasive imaging at synchrotron facilities. Iterative tomographic reconstruction algorithms are often preferred for re…
At-Scale Sparse Deep Neural Network Inference with Efficient GPU Implementation
Mert Hidayetoglu, Carl Pearson, Vikram Sharma Mailthody +4
This paper presents GPU performance optimization and scaling results for inference models of the Sparse Deep Neural Network Challenge 2020. Demands for network quality have increas…