3 papers
cs.LG2025
Exploration by Random Reward Perturbation
Haozhe Ma, Guoji Fu, Zhengding Luo +2
We introduce Random Reward Perturbation (RRP), a novel exploration strategy for reinforcement learning (RL). Our theoretical analyses demonstrate that adding zero-mean noise to env…
cs.LG2025
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
Thanh Vinh Vo, Young Lee, Haozhe Ma +2
Hidden confounders that influence both states and actions can bias policy learning in reinforcement learning (RL), leading to suboptimal or non-generalizable behavior. Most RL algo…
cs.CV2024
Decoupled Prompt-Adapter Tuning for Continual Activity Recognition
Di Fu, Thanh Vinh Vo, Haozhe Ma +1
Action recognition technology plays a vital role in enhancing security through surveillance systems, enabling better patient monitoring in healthcare, providing in-depth performanc…