2 papers
cs.LG2026
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
Dipendra Misra, Aldo Pacchiano, Ta-Chung Chi +1
We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by u…
cs.CL2025
A State-of-the-Art SQL Reasoning Model using RLVR
Alnur Ali, Ashutosh Baheti, Jonathan Chang +13
Developing custom reasoning models via Reinforcement Learning (RL) that can incorporate organization-specific knowledge has great potential to address problems faced by enterprise…