From the 1 of 1 linked paper with an AI index.
1 paper
Simone Drago, Marco Mussi, Leonardo Bianconi +1
The paper extends preference‑based reinforcement learning by allowing human experts to label trajectory pairs as incomparable, and introduces a Bradley‑Terry‑inspired rationality m…