2 citations · 3 across the 11 of their papers we have counts for
7 papers · 1 filter
Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation
Tung Tran, Viet Bao Mai, Hoang Ta +1
Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce…
Power Mean Estimation in Stochastic Continuous Monte Carlo Tree Search
Tuan Dam
Monte Carlo Tree Search (MCTS) has demonstrated success in online planning for deterministic environments, yet significant challenges remain in adapting it to stochastic Markov Dec…
Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing
Tuan Dam
Planning with a generative model aims to estimate the value of a state using as few simulator calls as possible. SmoothCruiser achieves problem-independent complexity $\widetilde O…
Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning
Hung Pham, Tuan Dam
Prioritized Sweeping (PS) accelerates model-based reinforcement learning by selecting backups according to Bellman residual magnitude. In nonstationary reward settings, however, th…
Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments
Khang Luong, Nam Nguyen, Hoang Ta +2
We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration process is dominated by stochast…
Conservation Laws for Modern Neural Architectures
Viet-Hoang Tran, Vinh Khanh Bui, Tan Lai Ngoc +3
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. Whi…