From the 1 of 1 linked paper with an AI index.
1 paper
Zhouchonghao Wu, Raymond Song, Vedant Mundheda +3
The paper introduces TADPO, a policy‑gradient method that extends PPO with teacher‑student guidance, and uses it in a vision‑based end‑to‑end reinforcement learning system for high…