9 citations · 9 across the 6 of their papers we have counts for
14 papers
DiffusionGemma Technical Report
DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at…
PDLP: A Practical First-Order Method for Large-Scale Linear Programming
David Applegate, Mateo Díaz, Oliver Hinder +4
We present PDLP, a practical first-order method for linear programming (LP) designed to solve large-scale LP problems. PDLP is based on the primal-dual hybrid gradient (PDHG) metho…
Probabilistic Inference in Reinforcement Learning Done Right
Jean Tarbouriech, Tor Lattimore, Brendan O'Donoghue
A popular perspective in Reinforcement learning (RL) casts the problem as probabilistic inference on a graphical model of the Markov decision process (MDP). The core object of stud…
On the connection between Bregman divergence and value in regularized Markov decision processes
Brendan O'Donoghue
In this short note we derive a relationship between the Bregman divergence from the current policy to the optimal policy and the suboptimality of the current value function in a re…
Variational Bayesian Optimistic Sampling
Brendan O'Donoghue, Tor Lattimore
We consider online sequential decision problems where an agent must balance exploration and exploitation. We derive a set of Bayesian `optimistic' policies which, in the stochastic…
Solving Mixed Integer Programs Using Neural Networks
Vinod Nair, Sergey Bartunov, Felix Gimeno +16
Mixed Integer Programming (MIP) solvers rely on an array of sophisticated heuristics developed with decades of research to solve large-scale MIP instances encountered in practice.…