papers

Publications (15)

cs.MS2022

Batched Second-Order Adjoint Sensitivity for Reduced Space Methods

François Pacaud, Michel Schanen, Daniel Adrian Maldonado +4

This paper presents an efficient method for extracting the second-order sensitivities from a system of implicit nonlinear equations on upcoming graphical processing units (GPU) dom…

cs.DC2023

Automated Translation and Accelerated Solving of Differential Equations on Multiple GPU Platforms

Utkarsh Utkarsh, Valentin Churavy, Yingbo Ma +8

We demonstrate a high-performance vendor-agnostic method for massively parallel solving of ensembles of ordinary differential equations (ODEs) and stochastic differential equations…

math.NA2026

Massively parallel numerical simulations with Julia

Simon Candelaresi, Benedict Geihe, Marco Artiano +6

The paper evaluates how the Julia programming language performs for large-scale, parallel computational fluid dynamics simulations, comparing it to a Fortran implementation and dem…

#parallel computing#numerical simulation#julia language#computational fluid dynamics
cs.DC2022

Bridging HPC Communities through the Julia Programming Language

Valentin Churavy, William F Godoy, Carsten Bauer +9

The Julia programming language has evolved into a modern alternative to fill existing gaps in scientific computing and data science applications. Julia leverages a unified and coor…

cs.DC2025

Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision

Evelyne Ringoot, Rabab Alomairy, Valentin Churavy +1

This paper presents a portable, GPU-accelerated implementation of a QR-based singular value computation algorithm in Julia. The singular value ecomposition (SVD) is a fundamental n…

math.NA2026

Automatic differentiation for performing the Cauchy-Kovalevskaya procedure in Lax-Wendroff type discretizations

Arpit Babbar, Valentin Churavy, Michael Schlottke-Lakemper +1

Lax-Wendroff methods combined with discontinuous Galerkin/flux reconstruction spatial discretization provide a high-order, single-stage, quadrature-free method for solving hyperbol…

astro-ph.IM2025

Hierarchical Interferometric Bayesian Imaging

Paul Tiede, William Moses, Valentin Churavy +4

Very long baseline interferometry (VLBI) achieves the highest angular resolution in astronomy. VLBI measures corrupted Fourier components, known as visibilities. Reconstructing on-…

physics.ao-ph2021

Large-eddy simulations with ClimateMachine: a new open-source code for atmospheric simulations on GPUs and CPUs

Akshay Sridhar, Yassine Tissaoui, Simone Marras +11

We introduce ClimateMachine, a new open-source atmosphere modeling framework using the Julia language to be performance portable on central processing units (CPUs) and graphics pro…

cs.DC2023

Evaluating performance and portability of high-level programming models: Julia, Python/Numba, and Kokkos on exascale nodes

William F. Godoy, Pedro Valero-Lara, T. Elise Dettling +7

We explore the performance and portability of the high-level programming models: the LLVM-based Julia and Python/Numba, and Kokkos on high-performance computing (HPC) nodes: AMD Ep…

physics.comp-ph2019

Highly-scalable, physics-informed GANs for learning solutions of stochastic PDEs

Liu Yang, Sean Treichler, Thorsten Kurth +8

Uncertainty quantification for forward and inverse problems is a central challenge across physical and biomedical disciplines. We address this challenge for the problem of modeling…

cs.DC2022

Bring the BitCODE -- Moving Compute and Data in Distributed Heterogeneous Systems

Wenbin Lu, Luis E. Peña, Pavel Shamis +3

In this paper, we present a framework for moving compute and data between processing elements in a distributed heterogeneous system. The implementation of the framework is based on…

cs.MS2020

Instead of Rewriting Foreign Code for Machine Learning, Automatically Synthesize Fast Gradients

William S. Moses, Valentin Churavy

Applying differentiable programming techniques and machine learning algorithms to foreign programs requires developers to either rewrite their code in a machine learning framework,…

cs.MS2018

Dynamic Automatic Differentiation of GPU Broadcast Kernels

Jarrett Revels, Tim Besard, Valentin Churavy +2

We show how forward-mode automatic differentiation (AD) can be employed within larger reverse-mode computations to dynamically differentiate broadcast operations in a GPU-friendly…

cs.DC2022

Productivity meets Performance: Julia on A64FX

Mosè Giordano, Milan Klöwer, Valentin Churavy

The Fujitsu A64FX ARM-based processor is used in supercomputers such as Fugaku in Japan and Isambard 2 in the UK and provides an interesting combination of hardware features such a…

physics.ao-ph2024

Oceananigans.jl: A Julia library that achieves breakthrough resolution, memory and energy efficiency in global ocean simulations

Simone Silvestri, Gregory L. Wagner, Christopher Hill +10

Climate models must simulate hundreds of future scenarios for hundreds of years at coarse resolutions, and a handful of high-resolution decadal simulations to resolve localized ext…