most citedQ-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

1 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Zero-Overhead Introspection for Adaptive Test-Time Compute

Rohin Manvi, Joey Hong, Tim Seyde +3

Large language models excel at reasoning but lack key aspects of introspection, including anticipating their own success and the computation required to achieve it. Humans use real…

cs.LG20251 cited

Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space

Joey Hong, Kang Liu, Zhan Ling +2

Large language model (LLM) agents -- LLMs that dynamically interact with an environment over long horizons -- have become an increasingly important area of research, enabling autom…

cs.CL2025

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

Joey Hong, Anca Dragan, Sergey Levine

Large language models (LLMs) excel in tasks like question answering and dialogue, but complex tasks requiring interaction, such as negotiation and persuasion, require additional lo…

cs.LG20241 cited

Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Joey Hong, Anca Dragan, Sergey Levine

Value-based reinforcement learning (RL) can in principle learn effective policies for a wide range of multi-turn problems, from games to dialogue to robotic control, including via…

cs.LG2024

Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations

Joey Hong, Jessica Lin, Anca Dragan +1

Recent progress on large language models (LLMs) has enabled dialogue agents to generate highly naturalistic and plausible text. However, current LLM language generation focuses on…