From the 1 of 1 linked paper with an AI index.
1 paper
Ilias Kazantzidis, Timothy J. Norman, Yali Du +1
The paper introduces DROPJ, a human‑in‑the‑loop approach that uses preferences and justifications collected from users interacting with a learned world model to train a reward mode…