1 paper
Suraj Yadav, Siddharth Yadav, Parth Goyal
Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this clai…