1 paper
Vatsal Ananthula, Adarsh Kumarappan
Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's internal reasoning. We propose verifier-cou…