5 papers
Performance Optimization and Comparative Analysis of Generative AI Models on Advanced Accelerators
Amitash Nanda, Javier Hernandez Nicolau, Madhusudan Gujral +3
Generative AI models, such as Large Language Models (LLMs) and diffusion models, have demonstrated impressive performance across a wide range of tasks. Despite these advances, depl…
The SDSC Satellite Reverse Proxy Service for Launching Secure Jupyter Notebooks on High-Performance Computing Systems
Mary P Thomas, Martin Kandes, James McDougall +4
Using Jupyter notebooks in an HPC environment exposes a system and its users to several security risks. The Satellite Proxy Service, developed at SDSC, addresses many of these secu…
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
Mohammad Firas Sada, John J. Graham, Elham E Khoda +7
This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throug…
Real-Time In-Network Machine Learning on P4-Programmable FPGA SmartNICs with Fixed-Point Arithmetic and Taylor
Mohammad Firas Sada, John J. Graham, Mahidhar Tatineni +3
As machine learning (ML) applications become integral to modern network operations, there is an increasing demand for network programmability that enables low-latency ML inference…
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
Derek Weitzel, Ashton Graves, Sam Albin +9
The National Research Platform (NRP) represents a distributed, multi-tenant Kubernetes-based cyberinfrastructure designed to facilitate collaborative scientific computing. Spanning…