1 paper
Subhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari +4
Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations,…