1 paper
Meng Cai, Lars Kulik, Farhana Choudhury
When language models use test-time sampling, they generate multiple reasoning trajectories and select an answer by majority vote. We show that these trajectories are not independen…