5 papers
An Alternating Approach to Approximate Dynamic Programming
Di Zhang
In this paper, we give a new approximate dynamic programming (ADP) method to solve large-scale Markov decision programming (MDP) problem. In comparison with many classic ADP method…
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
Di Zhang, Yihang Zhang
Stochastic gradient-based descent (SGD), have long been central to training large language models (LLMs). However, their effectiveness is increasingly being questioned, particularl…
Asymmetric Graph Error Control with Low Complexity in Causal Bandits
Chen Peng, Di Zhang, Urbashi Mitra
In this paper, the causal bandit problem is investigated, with the objective of maximizing the long-term reward by selecting an optimal sequence of interventions on nodes in an unk…
A Stochastic Conjugate Subgradient Algorithm for Two-stage Stochastic Programming
Di Zhang, Suvrajeet Sen
Stochastic Optimization is a cornerstone of operations research, providing a framework to solve optimization problems under uncertainty. Despite the development of numerous algorit…
An Adaptive Sampling-based Progressive Hedging Algorithm for Stochastic Programming
Di Zhang, Yihang Zhang, Suvrajeet Sen
The progressive hedging algorithm (PHA) is a cornerstone among algorithms for large-scale stochastic programming problems. However, its traditional implementation is hindered by so…