2 papers
cs.LG2026
Learning to Answer from Correct Demonstrations
Nirmit Joshi, Gene Li, Siddharth Bhandari +3
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at t…
cs.LG2025
Debiasing Reward Models by Representation Learning with Guarantees
Ignavier Ng, Patrick Blöbaum, Siddharth Bhandari +2
Recent alignment techniques, such as reinforcement learning from human feedback, have been widely adopted to align large language models with human preferences by learning and leve…