3 papers
cs.LG2025
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
Yoonho Lee, Joseph Boen, Chelsea Finn
We introduce \textit{Feedback Descent}, a framework that optimizes text artifacts -- prompts, code, and molecules -- through structured textual feedback, rather than relying solely…
cs.AI2025
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
Yuxiao Qu, Anikait Singh, Yoonho Lee +4
Reasoning requires going beyond pattern matching or memorization of solutions to identify and implement "algorithmic procedures" that can be used to deduce answers to hard problems…
cs.LG2024
Calibrating Language Models with Adaptive Temperature Scaling
Johnathan Xie, Annie S. Chen, Yoonho Lee +2
The effectiveness of large language models (LLMs) is not only measured by their ability to generate accurate outputs but also by their calibration-how well their confidence scores…