reinforcement learning 2autonomous driving 1end-to-end driving 1off-road navigation 1policy gradient 1procedural simulation 1self-play 1self-play reinforcement learning 1sim-to-real transfer 1simulation 1traffic rule enforcement 1vision alignment 1
From the 3 of 3 linked papers with an AI index.
3 papers
cs.CV2026
TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations
Zikang Xiong, Weixin Li, Zhouchonghao Wu +6
The paper proposes a method to train end-to-end autonomous driving policies without expert demonstrations by pretraining a policy via self‑play in a fast vectorized simulator and t…
cs.LG2026
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Zhouchonghao Wu, Akshay Rangesh, Weixin Li +5
TerraZero is a procedural driving simulator that enables large-scale, zero‑demonstration self‑play reinforcement learning for autonomous driving, achieving high simulation speed an…
cs.RO2026
TADPO: Reinforcement Learning Goes Off-road
Zhouchonghao Wu, Raymond Song, Vedant Mundheda +3
The paper introduces TADPO, a policy‑gradient method that extends PPO with teacher‑student guidance, and uses it in a vision‑based end‑to‑end reinforcement learning system for high…