papers

Publications (43)

stat.ML2026

Tensor Train Diffusion: Leveraging Low-Rank Structures for High-Dimensional Score-Based Sampling

Robert Gruhlke, Julius Berner, David Sommer +1

Diffusion models offer a powerful framework for sampling from complex probability densities by learning to reverse a noising process. A common approach involves solving for the tim…

cs.LG2024

Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEs

Md Ashiqur Rahman, Robert Joseph George, Mogab Elleithy +9

Existing neural operator architectures face challenges when solving multiphysics problems with coupled partial differential equations (PDEs) due to complex geometries, interactions…

cs.LG2024

Dynamical Measure Transport and Neural PDE Solvers for Sampling

Jingtong Sun, Julius Berner, Lorenz Richter +4

The task of sampling from a probability density can be approached as transporting a tractable density function to the target, known as dynamical measure transport. In this work, we…

cs.LG2023

The Modern Mathematics of Deep Learning

Julius Berner, Philipp Grohs, Gitta Kutyniok +1

We describe the new field of mathematical analysis of deep learning. This field emerged around a list of research questions that were not answered within the classical framework of…

cs.LG2025

Underdamped Diffusion Bridges with Applications to Sampling

Denis Blessing, Julius Berner, Lorenz Richter +1

We provide a general framework for learning diffusion bridges that transport prior to target distributions. It includes existing diffusion models for generative modeling, but also…

quant-ph2026

Fourier Neural Operators for Learning Dynamics in Quantum Spin Systems

Freya Shah, Taylor L. Patti, Julius Berner +3

Fourier Neural Operators (FNOs) excel on tasks using functional data, such as those originating from partial differential equations. Such characteristics render them an effective a…

cs.CV2025

Robust Representation Consistency Model via Contrastive Denoising

Jiachen Lei, Julius Berner, Jiongxiao Wang +5

Robustness is essential for deep neural networks, especially in security-sensitive applications. To this end, randomized smoothing provides theoretical guarantees for certifying ro…

cs.LG2023

How degenerate is the parametrization of neural networks with the ReLU activation function?

Julius Berner, Dennis Elbrächter, Philipp Grohs

Neural network training is usually accomplished by solving a non-convex optimization problem using stochastic gradient descent. Although one optimizes over the networks parameters,…

cs.LG2026

From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training

Julius Berner, Lorenz Richter, Marcin Sendera +2

We study the problem of training neural stochastic differential equations, or diffusion models, to sample from a Boltzmann distribution without access to target samples. Existing m…

cs.LG2024

Improved sampling via learned diffusions

Lorenz Richter, Julius Berner

Recently, a series of papers proposed deep learning-based approaches to sample from target distributions using controlled diffusion processes, being trained only on the unnormalize…

cs.CL2024

Large Language Models for Mathematicians

Simon Frieder, Julius Berner, Philipp Petersen +1

Large language models (LLMs) such as ChatGPT have received immense interest for their general-purpose language understanding and, in particular, their ability to generate high-qual…

cs.LG2025

Coarse Graining with Neural Operators for Simulating Chaotic Systems

Chuwei Wang, Julius Berner, Boris Bonev +6

Accurately predicting the long-term behavior of chaotic systems is crucial for various applications such as climate modeling. However, achieving such predictions typically requires…

cs.LG2020

Analysis of the Generalization Error: Empirical Risk Minimization over Deep Artificial Neural Networks Overcomes the Curse of Dimensionality in the Numerical Approximation of Black-Scholes Partial Differential Equations

Julius Berner, Philipp Grohs, Arnulf Jentzen

The development of new classification and regression algorithms based on empirical risk minimization (ERM) over deep neural network hypothesis classes, coined deep learning, revolu…

cs.LG2026

Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference

Denis Blessing, Julius Berner, Lorenz Richter +4

Solving stochastic optimal control problems with quadratic control costs can be viewed as approximating a target path space measure, e.g. via gradient-based optimization. In practi…

cs.CV2026

There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training

Jiachen Lei, Keli Liu, Julius Berner +4

Pixel-space generative models are often more difficult to train and generally underperform compared to their latent-space counterparts, leaving a persistent performance and efficie…

cs.LG2024

Solving Poisson Equations using Neural Walk-on-Spheres

Hong Chul Nam, Julius Berner, Anima Anandkumar

We propose Neural Walk-on-Spheres (NWoS), a novel neural PDE solver for the efficient solution of high-dimensional Poisson equations. Leveraging stochastic representations and Walk…

cs.CV2026

Parallel Decoding Distillation for Fast Image and Video Generation

Neta Shaul, Chao Liu, Arash Vahdat +1

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavi…

physics.ao-ph2026

Towards accurate extreme event likelihoods from diffusion model climate emulators

Peter Manshausen, Noah Brenowitz, Julius Berner +2

ML climate model emulators are useful for scenario planning and adaptation, allowing for cost-efficient experimentation. Recently, the diffusion model Climate in a Bottle (cBottle)…

cs.LG2026

Decoupled Diffusion Sampling for Inverse Problems on Function Spaces

Thomas Y. L. Lin, Jiachen Yao, Lufang Chiang +2

We propose a data-efficient, physics-aware generative framework in function space for inverse PDE problems. Existing plug-and-play diffusion posterior samplers represent physics im…

cs.CV2026

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Xinyin Ma, Julius Berner, Chao Liu +3

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional d…

cs.LG2026

Operator Learning Using Weak Supervision from Walk-on-Spheres

Hrishikesh Viswanath, Hong Chul Nam, Xi Deng +3

Training neural PDE solvers is often bottlenecked by expensive data generation or unstable physics-informed neural network (PINN) involving challenging optimization landscapes due…

cs.LG2025

Principled Approaches for Extending Neural Architectures to Function Spaces for Operator Learning

Julius Berner, Miguel Liu-Schiaffini, Jean Kossaifi +4

A wide range of scientific problems, such as those described by continuous-time dynamical systems and partial differential equations (PDEs), are naturally formulated on function sp…

cs.LG2025

Enabling Automatic Differentiation with Mollified Graph Neural Operators

Ryan Y. Lin, Julius Berner, Valentin Duruisseaux +5

Physics-informed neural operators offer a powerful framework for learning solution operators of partial differential equations (PDEs) by combining data and physics losses. However,…

cs.CV2026

Mode Seeking meets Mean Seeking for Fast Long Video Generation

Shengqu Cai, Weili Nie, Chao Liu +8

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to…

cs.LG2024

DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training

Zhongkai Hao, Chang Su, Songming Liu +6

Pre-training has been investigated to improve the efficiency and performance of training neural operators in data-scarce settings. However, it is largely in its infancy due to the…

cs.LG2026

A Library for Learning Neural Operators

Jean Kossaifi, Nikola Kovachki, Zongyi Li +8

We present NeuralOperator, an open-source Python library for operator learning. Neural operators generalize neural networks to maps between function spaces instead of finite-dimens…

cs.LG2026

Guided Diffusion Sampling on Function Spaces with Applications to PDEs

Jiachen Yao, Abbas Mammadov, Julius Berner +4

We propose a general framework for conditional sampling in PDE-based inverse problems, targeting the recovery of whole solutions from extremely sparse or noisy measurements. This i…

cs.LG2019

Towards a regularity theory for ReLU networks -- chain rule and global error estimates

Julius Berner, Dennis Elbrächter, Philipp Grohs +1

Although for neural networks with locally Lipschitz continuous activation functions the classical derivative exists almost everywhere, the standard chain rule is in general not app…

cs.LG2024

An optimal control perspective on diffusion-based generative modeling

Julius Berner, Lorenz Richter, Karen Ullrich

We establish a connection between stochastic optimal control and generative models based on stochastic differential equations (SDEs), such as recently developed diffusion probabili…

cs.LG2024

Neural Operators with Localized Integral and Differential Kernels

Miguel Liu-Schiaffini, Julius Berner, Boris Bonev +3

Neural operators learn mappings between function spaces, which is practical for learning solution operators of PDEs and other scientific modeling applications. Among them, the Four…

cs.LG2026

Self-Supervised Learning via Flow-Guided Neural Operator on Time-Series Data

Duy Nguyen, Jiachen Yao, Jiayun Wang +2

Self-supervised learning (SSL) is a powerful paradigm for learning from unlabeled time-series data. However, popular methods such as masked autoencoders (MAEs) rely on reconstructi…

stat.ML2025

Sequential Controlled Langevin Diffusions

Junhua Chen, Lorenz Richter, Julius Berner +3

An effective approach for sampling from unnormalized densities is based on the idea of gradually transporting samples from an easy prior to the complicated target distribution. Two…

cs.LG2023

Learning ReLU networks to high uniform accuracy is intractable

Julius Berner, Philipp Grohs, Felix Voigtlaender

Statistical learning theory provides bounds on the necessary number of training samples needed to reach a prescribed accuracy in a learning problem formulated over a given target c…

cs.CV2026

Transition Matching Distillation for Fast Video Generation

Weili Nie, Julius Berner, Nanye Ma +3

Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to…

cs.LG2025

Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning

Simon Frieder, Jonas Bayer, Sam Looi +13

The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…

cs.LG2020

Numerically Solving Parametric Families of High-Dimensional Kolmogorov Partial Differential Equations via Deep Learning

Julius Berner, Markus Dablander, Philipp Grohs

We present a deep learning algorithm for the numerical solution of parametric families of high-dimensional linear Kolmogorov partial differential equations (PDEs). Our method is ba…

cs.LG2025

Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing

Bingliang Zhang, Wenda Chu, Julius Berner +3

Diffusion models have recently achieved success in solving Bayesian inverse problems with learned data priors. Current methods build on top of the diffusion sampling process, where…

eess.IV2026

Physics-Aware Neural Operators for Direct Inversion in 3D Photoacoustic Tomography

Jiayun Wang, Yousuf Aborahama, Arya Khokhar +10

Learning physics-constrained inverse operators-rather than post-processing physics-based reconstructions-is a broadly applicable strategy for problems with expensive forward models…

cs.CV2026

Variational Flow Maps: Make Some Noise for One-Step Conditional Generation

Abbas Mammadov, So Takao, Bohan Chen +4

Flow maps enable high-quality image generation in a single forward pass. However, unlike iterative diffusion models, their lack of an explicit sampling trajectory impedes incorpora…

cs.LG2022

Robust SDE-Based Variational Formulations for Solving Linear PDEs via Deep Learning

Lorenz Richter, Julius Berner

The combination of Monte Carlo methods and deep learning has recently led to efficient algorithms for solving partial differential equations (PDEs) in high dimensions. Related lear…

cs.LG2026

Bridge Matching Sampler: Scalable Sampling via Generalized Fixed-Point Diffusion Matching

Denis Blessing, Lorenz Richter, Julius Berner +2

Sampling from unnormalized densities using diffusion models has emerged as a powerful paradigm. However, while recent approaches that use least-squares `matching' objectives have i…

cs.LG2023

Mathematical Capabilities of ChatGPT

Simon Frieder, Luca Pinchetti, Alexis Chevalier +5

We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets,…

cs.LG2026

Learning Lagrangian Interaction Dynamics with Sampling-Based Model Order Reduction

Hrishikesh Viswanath, Yue Chang, Aleksey Panas +3

Simulating physical systems governed by Lagrangian dynamics often entails solving partial differential equations (PDEs) over high-resolution spatial domains, leading to significant…