1.3k citations · 2.5k across the 155 of their papers we have counts for
20 papers · 1 filter
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Hao Li, Shamit Lal, Zhiheng Li +9
We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…
Self-Refining Diffusion Samplers: Enabling Parallelization via Parareal Iterations
Nikil Roashan Selvam, Amil Merchant, Stefano Ermon
In diffusion models, samples are generated through an iterative refinement process, requiring hundreds of sequential model evaluations. Several recent methods have introduced appro…
Non-Myopic Multi-Objective Bayesian Optimization
Syrine Belakaria, Alaleh Ahmadianshalchi, Barbara Engelhardt +2
We consider the problem of finite-horizon sequential experimental design to solve multi-objective optimization (MOO) of expensive black-box objective functions. This problem arises…
Convolutional Differentiable Logic Gate Networks
Felix Petersen, Hilde Kuehne, Christian Borgelt +2
With the increasing inference cost of machine learning models, there is a growing interest in models with fast and efficient inference. Recently, an approach for learning logic gat…
TrAct: Making First-layer Pre-Activations Trainable
Felix Petersen, Christian Borgelt, Stefano Ermon
We consider the training of the first layer of vision models and notice the clear relationship between pixel values and gradient update magnitudes: the gradients arriving at the we…
Newton Losses: Using Curvature Information for Learning with Differentiable Algorithms
Felix Petersen, Christian Borgelt, Tobias Sutter +3
When training neural networks with custom objectives, such as ranking losses and shortest-path losses, a common problem is that they are, per se, non-differentiable. A popular appr…