6 citations · 6 across the 5 of their papers we have counts for
5 papers · 1 filter
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7
We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…
Mitigating Downstream Model Risks via Model Provenance
Keyu Wang, Abdullah Norozi Iranzad, Scott Schaffter +3
Research and industry are rapidly advancing the innovation and adoption of foundation model-based systems, yet the tools for managing these models have not kept pace. Understanding…
On the Privacy of Selection Mechanisms with Gaussian Noise
Jonathan Lebensold, Doina Precup, Borja Balle
Report Noisy Max and Above Threshold are two classical differentially private (DP) selection mechanisms. Their output is obtained by adding noise to a sequence of low-sensitivity q…
Actor Critic with Differentially Private Critic
Jonathan Lebensold, William Hamilton, Borja Balle +1
Reinforcement learning algorithms are known to be sample inefficient, and often performance on one task can be substantially improved by leveraging information (e.g., via pre-train…
Neural Transfer Learning for Cry-based Diagnosis of Perinatal Asphyxia
Charles C. Onu, Jonathan Lebensold, William L. Hamilton +1
Despite continuing medical advances, the rate of newborn morbidity and mortality globally remains high, with over 6 million casualties every year. The prediction of pathologies aff…