activity
20242026
collaborators

6 papers

cs.LG2026

Auto-exploration for online reinforcement learning

Caleb Ju, Guanghui Lan

The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and action discounte…

cs.LG2026

Strongly-polynomial time and validation analysis of policy gradient methods

Caleb Ju, Guanghui Lan

This paper proposes a novel termination criterion, termed the advantage gap function, for finite state and action Markov decision processes (MDP) and reinforcement learning (RL). B…

cs.LG2026

Actor-Accelerated Policy Dual Averaging for Reinforcement Learning in Continuous Action Spaces

Ji Gao, Caleb Ju, Guanghui Lan +1

Policy Dual Averaging (PDA) offers a principled Policy Mirror Descent (PMD) framework that more naturally admits value function approximation than standard PMD, enabling the use of…

math.NA2025

Preconditioning via Randomized Range Deflation (RandRAND)

Oleg Balabanov, Caleb Ju, Kaiwen He +2

We introduce RandRAND, a new class of randomized preconditioning methods for large-scale linear systems. RandRAND deflates the spectrum via efficient orthogonal projections onto ra…

math.OC2025

A model-free first-order method for linear quadratic regulator with sampling complexity

Caleb Ju, Georgios Kotsalis, Guanghui Lan

We consider the classic stochastic linear quadratic regulator (LQR) problem under an infinite horizon average stage cost. By leveraging recent policy gradient methods from reinforc…

cs.LG2024

Learning a local trading strategy: deep reinforcement learning for grid-scale renewable energy integration

Caleb Ju, Constance Crozier

Variable renewable generation increases the challenge of balancing power supply and demand. Grid-scale batteries co-located with generation can help mitigate this misalignment. Thi…