Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
THOUGHTSCULPT: Reasoning with Intermediate Revision and Search
Yizhou Chi, Kevin Yang, Dan Klein
We present THOUGHTSCULPT, a general reasoning and search method for tasks with outputs that can be decomposed into components. THOUGHTSCULPT explores a search tree of potential sol…
cs.CL2024
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
Kevin Yang, Dan Klein, Asli Celikyilmaz +2
We propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more h…