3 papers
cs.LG2025
A Critical Look At Tokenwise Reward-Guided Text Generation
Ahmad Rashid, Ruotian Wu, Julia Grosse +2
Large language models (LLMs) can be improved by aligning with human preferences through fine-tuning -- the so-called reinforcement learning from human feedback (RLHF). However, the…
cs.LG2025
Uncertainty-Guided Likelihood Tree Search
Julia Grosse, Ruotian Wu, Ahmad Rashid +4
Tree search is a fundamental tool for planning, as many sequential decision-making problems can be framed as searching over tree-structured spaces. We propose an uncertainty-guided…
cs.LG2025
Towards Cost-Effective Reward Guided Text Generation
Ahmad Rashid, Ruotian Wu, Rongqi Fan +3
Reward-guided text generation (RGTG) has emerged as a viable alternative to offline reinforcement learning from human feedback (RLHF). RGTG methods can align baseline language mode…