19 citations · 44 across the 13 of their papers we have counts for
1 paper · 1 filter
Alejandro Escontrela, Xue Bin Peng, Wenhao Yu +4
Training a high-dimensional simulated agent with an under-specified reward function often leads the agent to learn physically infeasible strategies that are ineffective when deploy…