1 paper · 1 filter
Ebenezer Gelo, Geraud Nangue Tasse, Steven James +1
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first u…