25 citations · 34 across the 5 of their papers we have counts for
9 papers
Two-Stage Neural Contextual Bandits for Personalised News Recommendation
Mengyan Zhang, Thanh Nguyen-Tang, Fangzhao Wu +3
We consider the problem of personalised news recommendation where each user consumes news in a sequential fashion. Existing personalised news recommendation methods focus on exploi…
On Practical Reinforcement Learning: Provable Robustness, Scalability, and Statistical Efficiency
Thanh Nguyen-Tang
This thesis rigorously studies fundamental reinforcement learning (RL) methods in modern practical considerations, including robust RL, distributional RL, and offline RL with neura…
Offline Neural Contextual Bandits: Pessimism, Optimization and Generalization
Thanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen +1
Offline policy learning (OPL) leverages existing data collected a priori for policy optimization without any active exploration. Despite the prevalence and recent interest in this…
Combining Online Learning and Offline Learning for Contextual Bandits with Deficient Support
Hung Tran-The, Sunil Gupta, Thanh Nguyen-Tang +2
We address policy learning with logged data in contextual bandits. Current offline-policy learning algorithms are mostly based on inverse propensity score (IPS) weighting requiring…
Sample Complexity of Offline Reinforcement Learning with Deep ReLU Networks
Thanh Nguyen-Tang, Sunil Gupta, Hung Tran-The +1
Offline reinforcement learning (RL) leverages previously collected data for policy optimization without any further active exploration. Despite the recent interest in this problem,…
Distributional Reinforcement Learning via Moment Matching
Thanh Tang Nguyen, Sunil Gupta, Svetha Venkatesh
We consider the problem of learning a set of probability distributions from the empirical Bellman dynamics in distributional reinforcement learning (RL), a class of state-of-the-ar…