14 citations · 28 across the 4 of their papers we have counts for
1 paper · 1 filter
Alejandro Escontrela, Xue Bin Peng, Wenhao Yu +4
Training a high-dimensional simulated agent with an under-specified reward function often leads the agent to learn physically infeasible strategies that are ineffective when deploy…