2 papers
cs.LG2025
Transformer Reconstructed with Dynamic Value Attention
Xiaowei Wang
Since transformer was firstly published in 2017, several works have been proposed to optimize it. However, the major structure of transformer remains unchanged, ignoring one of its…
cs.LG2021
Delayed Rewards Calibration via Reward Empirical Sufficiency
Yixuan Liu, Hu Wang, Xiaowei Wang +3
Appropriate credit assignment for delay rewards is a fundamental challenge for reinforcement learning. To tackle this problem, we introduce a delay reward calibration paradigm insp…