Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Making Expert Reasoning Learnable with Self-Distillation
Ethan Mendes, Jungsoo Park, Alan Ritter
Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct solution to be reinforced or the existence o…
cs.LG2026
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
Jungsoo Park, Hyungjoo Chae, Ethan Mendes +4
Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floati…
cs.LG2025
Language Models can Self-Improve at State-Value Estimation for Better Search
Ethan Mendes, Alan Ritter
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, particularly in interactive domains such as web tasks. We i…